Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

61–70 of 374 posts

Re: An update on Wayback Machine access

#61
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…

What if they charged money? Is it something you'd pay for?

Re: An update on Wayback Machine access

#63
Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

If you got some money to spare, consider donating to them. They need it.

Re: An update on Wayback Machine access

#64
post #48
post #2

Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon. It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs. I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

> Why not just offer a paid endpoint for the crawlers? Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.

Re: An update on Wayback Machine access

#65
post #49

Earlier quoted context omitted.

Just yesterday from my one of my sessions with Sol: > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

As it should. Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

Re: An update on Wayback Machine access

#66
post #64
post #48

Earlier quoted context omitted.

> Why not just offer a paid endpoint for the crawlers? Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.

Exactly, it's a question for the lawyers to sort out.

Re: An update on Wayback Machine access

#67
post #13

Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.

Try making a vpn via digital ocean for example and you'll see similar patterns.

Re: An update on Wayback Machine access

#69
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

Do they offer bulk torrent downloads as an alternative?

Re: An update on Wayback Machine access

#70
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

> I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.

> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.

Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.

Post reply on HN