Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

221–230 of 374 posts

Re: An update on Wayback Machine access

#221
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I say name and shame!

Re: An update on Wayback Machine access

#222
post #111

Earlier quoted context omitted.

Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)

You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason . The end result is exactly the same.

It is different, precisely because the end result is not the same - one broadly benefits the public while the other doesn't.

Substitute almost any disruptive public service to see the issue with your line of reasoning. For example - you seem to think that [ bulldozing private property ] to "construct an emergency fire break" is somehow different than [ bulldozing private property ] for any other reason.

Never mind that the sort of scraping being objected to is actually harmful to service health while what the wayback machine does is almost entirely unnoticeable.

Re: An update on Wayback Machine access

#223

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.

> Shouldn't the solution be to gate bulk access for automated services for a price? Fine in theory but determined scrapers will use residential proxies in bulk.

These should be illegal unless users sign off on every fucking byte.

Re: An update on Wayback Machine access

#224
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

My use of wayback has skyrocketed recently due to anti-bot measures.

I often cannot get past captchas, and archive.org is one of the fallbacks I try.

However, archive.is, etc are more reliable.

I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.

They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.

Re: An update on Wayback Machine access

#225
post #13

Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

Are there any decent/reputable residential proxy companies? I’m pretty sure I’ll end up needing one occasionally, for those days when my residential IP has a poor reputation score.

Re: An update on Wayback Machine access

#227

Earlier quoted context omitted.

I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

From a few weeks ago: https://news.ycombinator.com/item?id=49500040 Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate. You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.

[dead]

Re: An update on Wayback Machine access

#228
post #167

Earlier quoted context omitted.

What is 'insane' here is the shear level of entitlement displayed here, including lumping a niche, free, volunteer supported service in with billion dollar, for profit corporations and demanding they pander to your inflated expectations. Wild.

It doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.

> details that could be captured automatically through web logs

You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that info?

> They created a problem

No, they're dealing with a problem, and compromised that some human users may unfortunately get blocked.

> and now users have to pay for the inconvenience

You don't have to anything. You can just not use them. They don't owe you their service.

Somebody is handing out free apple lollipops, they ran out, compromised on giving grape ones, and now you're complaining you're being forced to eat a grape one and you don't like grape. Don't eat it.

Re: An update on Wayback Machine access

#229

Earlier quoted context omitted.

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a…

How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

They have their own ASN, which I’ve explicitly allowed requests from.

https://www.peeringdb.com/asn/7941

Re: An update on Wayback Machine access

#230
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

Perhaps it would be sensible for the wayback machine to not show paywalled articles for the first, say, 3 months.
Post reply on HN