Web Archives and Big Data
An introduction to the process of Web archiving as well as to the characteristics of national Web archives by professor Niels Brügger, Aarhus University.
National and transnational Web archives have been established since the mid-1990s, and they continue to grow and thereby accumulate large amounts of data. Taking these massive collections of data as point of departure this article discuss the following two questions:
- To what extent can Web archives be considered big data?
- How does the nature of the archived Web in (trans)national Web archives impact the study of Web archives as big data?
A brief introduction to the process of Web archiving as well as to the characteristics of national Web archives is followed by a short summary of the considerations about big data put forward in the monograph Big Data: A revolution that will transform how we live, work, and think.
It is concluded, first, that Web archives may be considered big data in terms of size - they are, in fact, big - but not in the sense that to a large extent they are selections, everything that was on the Web is not necessarily in the Web archive. Second, Web archives may be considered big data since they are messy, but for the same reason it may be hard to correlate datasets in the Web archive.
With this conclusion as a stepping stone it is examined how scholars could study this specific type of big data, and the possible inclusion of more Web archives thus moving from big data to bigger data is debated.