Live data from Hacker News

To preserve their work journalists take archiving into their own hands

niemanlab.org

71–80 of 94 posts

Re: To preserve their work journalists take archiving into their own hands

#71

Earlier quoted context omitted.

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

But wouldn't that be a big waste if everyone who references the same thing is then keeping a copy of it.

> big waste

Storage cost has fallen exponentially for decades, https://ourworldindata.org/data-insights/the-price-of-comput...

Re: To preserve their work journalists take archiving into their own hands

#72
post #32

Earlier quoted context omitted.

Two big issues with Archive.org are that 1. it's a single point of failure, they don't encourage mirror sites to emerge, and 2. they keep using the "brand" to fight unwinnable battles like hosting books they don't own online, risking the whole endeavor. I still appreciate it, but just imagine if it goes down due to a lawsuit. Now that Google no longer shows cached results, an entire historical record would be gone.

Its surprising that archive.org is the only such outfit I have encountered. Just like we have had libraries since ancient times, why are there so few digital libraries? There must be others, but nowhere near the number (or awareness) that we should have. Heck, existing paper-based libraries should probably each include a digital archiving department. Maybe this is already happening or already exists, and is trivial t…

Excellent question.

Local neighborhood libraries could have their own curated digital archive, as cache for fast local search, and archival backup for long-term resilience.

Re: To preserve their work journalists take archiving into their own hands

#73
Ever since the NYT legal case against OpenAI (pronounce: ClosedASI, not FossAGI; free as in your data for them, not free as in beer) there seems to be an underground current pulling into a riptide of closed information access on the web. Humorously enough the zimmit project has been quietly updating the living heck out of itself and awakening from a nearly 6-8 year slumber. The once simple format for making a mediawiki offline archive now is able to mirror any website complete with content such as video, pdf, or other files.

It feels a lot like the end of usenet or geocities, but this time without the incentive for the archivists to share their collections as openly. I am certain full scrapes of reddit and twitter exist, even with post API closure changes, but we will likely never see these leave large AI companies internal data holdings.

I have taken it upon myself to begin using the updated zimmit docker container to start archiving swaths of the 'useful web', meaning not just high quality language tokens, but high quality citations and knowledge built with sources that are not just links to other places online.

I started saving all my starred github repos into a folder and it came out just around 125gb of code.

I am terrified that in the very short term future a lot of this content will either become paywalled or the financial incentives of hosting large information repositories will increase past the point of current ad revenue based models as more powerful larger scraping operations seek to fill their petabytes while i try to prevent my few small TB of content i dont want to lose from slipping through my fingers.

If anyone actually cares deeply about content preservation, go and buy yourself a few 10+ TB external disks and grab a copy of zimmit and start pulling stuff. Put it on archive.org and tag it. So far the only zim files I see on archive.org are the ones publicly released by the kiwix team yet there is an entire wiki of wikis called wikiindex that remains almost completely unscraped. Fandom and Wikia are gigantic repositories of information and I fear they will close themselves up sooner than later, while many of the smaller info stores we have all come to take for granted as being "at our fingertips" will slowly slip away.

I first noticed the deep web deepening when things I used to be able to find on google were no longer showing up no matter how well I knew the content I was searching for, no matter the complex dorking i attempted using operators in the search bar, just like it had vanished. For a time bing was excellent at finding these "scrubbed" sites. Then duckduckgo entered the chat, and bing started to close itself down more. Bing was just a scrape of google, and google stopped being reliable, so downstream "search indexers" just became micro googles that were slightly out of date with slightly worse search accuracy, but those ghost pages were now being "anti-propagated" into these downstream indexers.

Yandex became and is still my preferred search engine when I actually need to find something online, especially when using operators to narrow wide pools.

I have found some rough edges with zimmit and I am planning on investigating and even submitting some PR upstream, but when an archive attempt takes 3 days to run before crashing and then wiping out the progress it has been hard to debug without the FOMO hitting that I should be spending the time getting what I can now before coming back to work on the code and get everything properly.

If any have the time to commit to the project and help make it more stable, perhaps work on some more fault recovery or failure continuation it would make archivists like me who are strapped for time very very happy.

Please go and make a dent in this, news is not the only part of the web i feel could be lost forever if we do not act to preserve it.

In 5 years time I see generic web searches being considered a legacy software and eventually decommissioned in favor of AI native conversational search (blow my brains out). I know for a fact all AI companies are doing massive data collection and structuring for graphrag style operations, my fear is that when its working well enough search will just vanish until a group of hobbyists make it available to us again.

Re: To preserve their work journalists take archiving into their own hands

#74
post #66

Earlier quoted context omitted.

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

certainly not GitHub

What would you recommend instead?

Re: To preserve their work journalists take archiving into their own hands

#75

Earlier quoted context omitted.

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

But wouldn't that be a big waste if everyone who references the same thing is then keeping a copy of it.

> everyone who references the same thing is then keeping a copy of it.

... and it would serve as a form of redundancy, imitating the fungible nature of physical media: In order to cite the latest, copied, manuscript (for example) you needed to own a physical copy. They existence of these has enabled survival of works that would otherwise have been lost, or even reconstruction through ecdotics.-

Re: To preserve their work journalists take archiving into their own hands

#77

Earlier quoted context omitted.

I wrote a book in 2010. It had a references section with links to about 100 websites. When I wrote the second edition only about five years later, 50% of those links no longer worked. What we're doing right now is borderline insane. We're putting all of this information on the web, but almost each individual bit of information is dependent on either a company or a human being keeping it online. It's inevitable that c…

> so almost all of the information that is online right now will just disappear in the next 80 years. > And we essentially only have one single entity that tries to retain that information. Will future ages find ours a dark age, a gap in their records, a void ... ... up until the point - if ever - where a sufficiently advanced solution for permanence is found and comes online?

> ... up until the point - if ever - where a sufficiently advanced solution for permanence is found and comes online?

Like the laser printer?

The cost of permanent, physical preservation is pennies. People just don't do it for most things. And it doesn't guarantee accessibility, which has hosting costs.

Re: To preserve their work journalists take archiving into their own hands

#78
post #32

Earlier quoted context omitted.

Archive.org is such a godsend.- The entire information 'substrate' of society is ephemeral , if digital, and none (at least not enough) seem to have noticed .-

Two big issues with Archive.org are that 1. it's a single point of failure, they don't encourage mirror sites to emerge, and 2. they keep using the "brand" to fight unwinnable battles like hosting books they don't own online, risking the whole endeavor. I still appreciate it, but just imagine if it goes down due to a lawsuit. Now that Google no longer shows cached results, an entire historical record would be gone.

> I still appreciate it, but just imagine if it goes down due to a lawsuit. Now that Google no longer shows cached results, an entire historical record would be gone.

Or somebody accidentally `rm -rf`'s an empty variable. Or The Big One hits San Fran. Or somebody in crisis breaks in with a crowbar, matchbook, and jug of gasoline.

They're a rather old-school shop. Own their own servers, all in one location I think. Bare metal admin stuff, and data's only mirrored across two disks per file IIRC. Keeps costs down. It's what makes the whole operation possible. But I also wonder sometimes.

Re: To preserve their work journalists take archiving into their own hands

#79
post #32

Earlier quoted context omitted.

Two big issues with Archive.org are that 1. it's a single point of failure, they don't encourage mirror sites to emerge, and 2. they keep using the "brand" to fight unwinnable battles like hosting books they don't own online, risking the whole endeavor. I still appreciate it, but just imagine if it goes down due to a lawsuit. Now that Google no longer shows cached results, an entire historical record would be gone.

Its surprising that archive.org is the only such outfit I have encountered. Just like we have had libraries since ancient times, why are there so few digital libraries? There must be others, but nowhere near the number (or awareness) that we should have. Heck, existing paper-based libraries should probably each include a digital archiving department. Maybe this is already happening or already exists, and is trivial t…

There are lots of web archiving projects out there:

https://en.wikipedia.org/wiki/List_of_Web_archiving_initiati...

But the web is large. And public sector or academic librarian teams tend to be small. The IA's the one that people have coalesced around.

Re: To preserve their work journalists take archiving into their own hands

#80
post #30

A nice social attack is to create an internet archive looking website call it archive.newtld and use it to create social proof of things you didn't actually do. "Oh yeah the Washington Post did a redesign but here are my past 10 posts which I saved in archive: link " In post truth internet, proving archives is going to be tough and unless there's some other form of verification it's going to be useless fast for "prov…

Could this be solved by digital signatures on web content? (Or, a way to store those)

If you know you might want to prove something's authenticity later, post the SHA256 somewhere now. E.G. Multiple social media sites. Large and trusted web archives. Cryptocurrency blockchains. The last one's a stronger proof, with lots of money making sure it stays immutable.

Or hash everything you produce/consume. Then hash the list of hashes, and post that.

Or alternatively, counter forgeries by capturing more data. For web, but all data in general. Sensor RAWs to prove an image isn't "AI". Browser and network stack RAM dumps to prove a website's authenticity. Etc. There's what, a couple dozen accelerometers, GPS's, LIDARs etc. on the latest iPhones?

Post reply on HN