Live data from Hacker News

News outlets are limiting the Internet Archive’s access to their journalism

niemanlab.org

101–110 of 130 posts

Re: News outlets are limiting the Internet Archive’s access to their journalism

#101
post #15

That's a real shame. I am involved with some history-related projects and the number of websites which go offline is huge, and the wayback machine is incredibly helpful for unearthing these dead sites. It is not hard to imagine a future in 50 years time where a huge percentage of this content is lost forever, or at best incredibly hard to find.

Unfortunately the IA itself has become way less usable as their aggressive anti-bot protection means that actually doing any kind of (manual) explorative research, as opposed to pulling an individual website that you already know the exact URL of, is more than likely to get you temp banned.

Re: News outlets are limiting the Internet Archive’s access to their journalism

#102
post #79

Earlier quoted context omitted.

This future is here already, policy makers have it locked up. Any person who remembers what microfiche is understands the magnitude of this problem of not having a trustworthy public record. If we extended public policy from the library era, the library of congress itself would be the Internet Archive.

> If we extended public policy from Similarly and tangentially, when the US Constitution was made in an era of horseback/carriages, it explicitly authorized the creation of a public national postal service (USPS). If we extended that older public policy with today's technological context, they would have authorized a national Internet Service Provider. (And, like with USPS, specialized private competitors would exist…

That's at least better than other countries who have essentially privatized most of their existing public infrastructure.

Re: News outlets are limiting the Internet Archive’s access to their journalism

#103

Earlier quoted context omitted.

Good idea, but only if the article can't be edited during that week. What's worth preserving is the version the audience actually read. Articles routinely get ninja-edited after publication, sometimes repeatedly. Changelogs should be mandatory but they're useless if we can't keep them honest.

Block public access not archival

The reason they're blocking archives is people can go to the archive, to bypass paywalls and avoid targeted adverts, instead of the news site. It's also to prevent AI scrapers harvesting articles.

Re: News outlets are limiting the Internet Archive’s access to their journalism

#104
post #99

There's an incredibly simple fix: block the archive for a week. No one is paying after a week, so you let the Archive access after that. I don't see why every news outlet doesn't just do this.

Do these major publications charge per article? They should, but they don't. So their whole sell is that in aggregate (so access to all, including old articles) they are worth paying monthly for. In which case archive is a major revenue slumper

How would archive not be a revenue drain if there was pay per read articles? I would think the incentive to try to find a free version would increase not decrease, especially for a wide class of articles that are basically, “I’m curious but not that curious” which in aggregate I might pay money for (they add value to my subscription) but individually feel wasteful (do I really want to pay to satisfy this curiosity?)

Re: News outlets are limiting the Internet Archive’s access to their journalism

#105

Newspapers are failing at an astounding rate. Archive.org is just a (poor) scapegoat for their inability to survive. This makes the point everyone else is making even more important - that those stories need to be archived before they are lost for all time. "Since the early 2000s, the U.S. has lost about 40% of its local newspapers and about 75% of the jobs in newspaper journalism, according to a 2025 report from the…

There is a future where AI companies start hiring their own reporters, and it might be sooner rather than later.

Based on the coverage it looks a lot like journalists and publications have already been bought and paid for.

Re: News outlets are limiting the Internet Archive’s access to their journalism

#106

Earlier quoted context omitted.

How about I2P?

All privacy is an illusion, the government can read your thoughts, your neighbors are secretly informants, etc etc. Tor is fine especially for onion sites. You just have to understand the limitations. (I2P is also good.)

[deleted]

Re: News outlets are limiting the Internet Archive’s access to their journalism

#107

There really should be a micropayments setup on the internet that's not advertising based. Let these models pay a nickel to read the article, covered by the multi trillion dollar AI blank check.

We are working on this and have a live solution, payments for bot fetching of articles, while letting humans and certain approved bots read for free. Check out: https://proofivy.com/blog/cryptoslate_x402_pay_per_article_i...

Re: News outlets are limiting the Internet Archive’s access to their journalism

#108
post #15

That's a real shame. I am involved with some history-related projects and the number of websites which go offline is huge, and the wayback machine is incredibly helpful for unearthing these dead sites. It is not hard to imagine a future in 50 years time where a huge percentage of this content is lost forever, or at best incredibly hard to find.

I gave a talk about this when I worked for The Archive. There was an article in Scientific American about how the average lifetime of a page on the net before it 404s is about 100 days. That article is offline now and we accessed it through the wayback machine.

My own last project before I left was to ingest records from crawl dumps from the defunct cuil.com website. About 200 TB of stuff that brought back 60 billion URLs.

The nature of the internet has changed and it's become an ephemeral place for many people where you just through things in and others mine it as "data".

Re: News outlets are limiting the Internet Archive’s access to their journalism

#109
My cynical view is that a lot of these outlets would have liked to block the archive anyway⁰ but didn't as it could look bad to do so, and AI scraping is a convenient excuse. Much like some (but far from all) of the recent job cuts that have been announced “due to AI”.

An even more cynical view is that the information on many local news sites in recent years isn't worth archiving anyway, it is largely generic rubbish filtered down from on high because these days most local outlets are owned by large national groups¹ that use them for little more than a place to insert adverts.

For actual local news, which those outlets do sometimes still carry, archiving personal blogs, event sites, and some social media content³, would be more useful than local news outlets. It is a shame that a lot of this has moved to platforms that are more difficult to archive (distraction media providers block the archive too, discord and similar services are more difficult to easily/meaningfully archive, heck searching for non-recent information on them that you know is there can be a pain, etc.)

--------

[0] For numerous reasons including some idea that it could affect their advertising revenue, that they don't want things which they correct [because they are actually wrong or because those higher up the ownership chain are happy with certain truths] are embarrassingly preserved in their original form, etc.

[1] Like Reach Plc or Newsquest¹ (which owns the most prominent local rag where I live) in the UK.

[2] Which is in turn owned by the US company USA Today Co.

[3] Though this is probably largely blocked from the archive too.

Re: News outlets are limiting the Internet Archive’s access to their journalism

#110
post #61

Or maybe, just maybe these news sites shouldn't be shipping 40MB JS bloated, ad infected websites. You're a news station just ship the words, make people pay for the images. This keeps bandwidth down for non payers, and foots the bill for those who do use the bandwidth. You pay for what you use, and reduce the overhead while you're at it.

Bandwidth is not their concerns.

Their concern is profit. Text only pages has little footprint both on the network and on storage. I was trying to be concise.
Post reply on HN