> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…
An update on Wayback Machine access
221–230 of 374 posts
Re: An update on Wayback Machine access
#222Earlier quoted context omitted.
Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason . The end result is exactly the same.
Substitute almost any disruptive public service to see the issue with your line of reasoning. For example - you seem to think that [ bulldozing private property ] to "construct an emergency fire break" is somehow different than [ bulldozing private property ] for any other reason.
Never mind that the sort of scraping being objected to is actually harmful to service health while what the wayback machine does is almost entirely unnoticeable.
Re: An update on Wayback Machine access
#223Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.
> Shouldn't the solution be to gate bulk access for automated services for a price? Fine in theory but determined scrapers will use residential proxies in bulk.
Re: An update on Wayback Machine access
#224> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…
I often cannot get past captchas, and archive.org is one of the fallbacks I try.
However, archive.is, etc are more reliable.
I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.
They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.
Re: An update on Wayback Machine access
#225Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
Re: An update on Wayback Machine access
#226Re: An update on Wayback Machine access
#227Earlier quoted context omitted.
I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...
From a few weeks ago: https://news.ycombinator.com/item?id=49500040 Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate. You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.
Re: An update on Wayback Machine access
#228Earlier quoted context omitted.
What is 'insane' here is the shear level of entitlement displayed here, including lumping a niche, free, volunteer supported service in with billion dollar, for profit corporations and demanding they pander to your inflated expectations. Wild.
It doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.
You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that info?
> They created a problem
No, they're dealing with a problem, and compromised that some human users may unfortunately get blocked.
> and now users have to pay for the inconvenience
You don't have to anything. You can just not use them. They don't owe you their service.
Somebody is handing out free apple lollipops, they ran out, compromised on giving grape ones, and now you're complaining you're being forced to eat a grape one and you don't like grape. Don't eat it.
Re: An update on Wayback Machine access
#229Earlier quoted context omitted.
I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a…
How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...
Re: An update on Wayback Machine access
#230> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…