A common misapprehension, even amongst those who have been engaged at foundationally-deep levels of the Web, is that
the World Wide Web is not an archival
mechanism but a publishing
mechanism.More specifically, it's an on-demand publishing mechanism, relating to media-based publishing (books, records, physical video media) much in the same was as the electrical grid (a transmission mechanism) does to fuels (a storage function).
It would be nice (in at least some regards) if publishing content at a specific addressable URL were a promise to 1) eternally provide that resource and 2) never change it. But there's no way to guarantee that this will be the case.
In the past, the means for achieving archival of information was:
- To specifically record that information in some form. As with, say, Plato's memorialising of Socrates's dialgogues.
- To create multiple copies of those recordings, so that loss of any one instance doesn't mean total loss.
- To define a refererencing or indexing system such that individual works can be identified unambiguously. (Or more specifically, with an acceptable level of ambiguity.)
- To develop a means of agreement as to what the canonical or recongised version of a work is, or absent that, of identifying more canonical forms or lineages. (The historiography of the history of philosophy is an interesting sub-field, with a good introductory treatment in Peter Adamson's History of Philosophy Without Any Gaps podcast.)
URLs of and by themselves address these needs poorly. At best they point to a location at which a document may have been available at a point in time. A URL plus a time-range (which is what the Internet Archive's Wayback Machine effectively delivers) is a much better approximation to what is needed. (The notion of time-bounded identifiers generally seems a useful one, as might be applied also to domain and user names, for example.)
As I see it, the goal of the archival web needs to address a number of points which Zittrain addresses only very indirectly:
- Identification of what really should be archived. Right To Be Forgotten exists for very well-founded reasons, and a world without forgetting (or with very capricious forgetting) is one form of hell.
- True document-centric identification. I've been thinking about this for a while (recent discussion here: https://news.ycombinator.com/item?id=27455520), and something that is based on the actual contents while being resilient to mild changes seems most optimal. Checksums, not so much, tuples or ngrams, possibly warmer.
- A number of archival institutions. The Internet Archive is certainly amongst these. Other libraries, if at all possible spread across multiple institutions and jurisdictions would be preferable. IA have been working with numerous academic institutions and the US Library of Congress, though IA itself still carries most of the burden.
This isn't a new problem. In particular, each time there's been an explosion in some new form of publishing, there's been a scramble by archivists to keep up. Denis Diderot, 18th century encyclopaedist, has an awesome quote about the information explosion drowning his own generation (the encyclopaedia was his technical solution to that problem).[1] Numerous elements of what we now accept as standard bibliographic elements (titles, authors, tables of contents, indices, references, citations, page numbers, paragraphs, inter-word spaces, ...) were invented to answer specific needs, not always of archivists as a principle focus, though often providing benefits to them. Cataloguing and classification systems likewise.
I appreciate Zittrain's message. He's crying over a lost cause and looking backwards, not to the future.
________________________________
Notes:
1. Diderot: https://www.historyofinformation.com/detail.php?entryid=2877