Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

351–360 of 374 posts

Re: An update on Wayback Machine access

#351
post #346

I think it's a matter of time before archive.org gets "bought" and dissappears. There should be government sponsored mirrors in many places of the world. The amount of data in archive.org is about 100PB. We're talking 10 racks of disks. I think archive.org should sell "archive as a service" for let's say $15mln. Half of that would be hardware cost and the deliverable could be 12 DC racks containing entire archive.org…

The archive is a non profit funded by various foundations and a a congressionally designated depository for U.S. Government documents. They may have a job selling it off without objections.

They should sell copies. I’m sure some foundation model company would happily hand them more than enough to establish a self sustaining foundation. Also, then there would be multiple copies. They could even give torrent access to libraries.

Re: An update on Wayback Machine access

#352
post #236

Earlier quoted context omitted.

Convenience as a higher order motivator than disgust at the bad behavior of archive.{today,ph,...} mentioned elsewhere, I think is the point of the comment to which you replied.

Convenience is by definition one step. Disgust requires research, analysis, and decision. The research alone is beyond most users' general practice. Why would anyone walk in the front door when they could hop a fence, pry open a window, and crawl in?

disgust is the simplest of emotions and so requires none of what you claim.

You're mistaking it with substantiated criticism.

Re: An update on Wayback Machine access

#353
I've been doing a bit of (very careful to be polite) scraping of Wayback to get archives of now-offline sites, so I hope this won't cause any significant issues for me. I did contact them beforehand to request a direct copy of the sites/networks required (which I believe they offered at some point) but unfortunately received no reply.

Re: An update on Wayback Machine access

#354
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I think the Wayback Machine offers an official API for bots to call.

Re: An update on Wayback Machine access

#355
post #351
post #346

Earlier quoted context omitted.

The archive is a non profit funded by various foundations and a a congressionally designated depository for U.S. Government documents. They may have a job selling it off without objections.

They should sell copies. I’m sure some foundation model company would happily hand them more than enough to establish a self sustaining foundation. Also, then there would be multiple copies. They could even give torrent access to libraries.

If they did this, many websites would immediately opt out or block them.

Re: An update on Wayback Machine access

#356
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

[dead]

Re: An update on Wayback Machine access

#357

Earlier quoted context omitted.

The other parts are even more trivial for someone who cares to work around.

"someone who cares" is the part that creates the firewall.

With modern AI tools, the amount you have to care is slight. Not much of a firewall. And the bad actors causing these scraping problems care a lot.

Re: An update on Wayback Machine access

#358

The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that. I knew something was up when a few alternative YouTube front-ends I use suddenly put up the…

[dead]

Re: An update on Wayback Machine access

#360

Earlier quoted context omitted.

You can't make other people forget things by declaring yourself ignorant of the facts.

I literally had never heard of this before. I don't check HN every single day. It's extremely reasonable to ask for a link, very easy to include one when making a claim, and attacking someone for asking for evidence is extremely anti-intellectual independent of the level of effort required. (n.b. that doesn't excuse the hostile way that they asked for proof - "Let's all just believe this baseless assertion shall we")

[dead]
Post reply on HN