Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

191–200 of 374 posts

Re: An update on Wayback Machine access

#191
post #118

The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.

How exactly does IPv6 "stop crawlers".

If anything, it will make it harder to block due to the vastness of the IPv6 address space.

Re: An update on Wayback Machine access

#192
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I wonder if the entire internet is going to slowly move behind logins and allow lists for specific trusted crawlers at some point. Open access doesn't seem sustainable. But I might just grumpy about spending another hour this week adjusting rules to prevent bots.

Rather than logins or regional filters, how about they just be a content provider to local libraries and perhaps use an app like Libby.

Re: An update on Wayback Machine access

#193
post #101

Earlier quoted context omitted.

Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine , it very much is a distinction with a difference.

Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access. So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

Internet Archive's traffic may not matter to you, but that's the main topic of this discussion, regardless of what you care or use website archival tools for.

Re: An update on Wayback Machine access

#194

Earlier quoted context omitted.

No matter what they did, you'd have a new excuse for why you won't pay.

Ah, the old ad hominem attack. How refreshing. But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but". It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

If you're donating, then you are evidently willing to pay without any "buts". So aren't you arguing against your own actions?

Re: An update on Wayback Machine access

#195

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access. The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gat…

I've been donating $5 to them monthly for I don't know how long. I've only recently bumped it up to $25. They're the heroes of the internet age

Re: An update on Wayback Machine access

#196
post #97

Earlier quoted context omitted.

> Asking users to email them with details of their OS, browser, IP address is just crazy. It's surely to serve as data to help tell humans apart from bots. > Changes made by IA shouldn't become my responsibility. They're a free service. It's ultimately not their responsibility to service you either.

Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request. IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

Who do you think the Internet Archive is? They are not Google, they have very low funding, very high expenses and are constantly under legal pressure.

The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.

Re: An update on Wayback Machine access

#197
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests.

However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.

Re: An update on Wayback Machine access

#198

Earlier quoted context omitted.

Ah, the old ad hominem attack. How refreshing. But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but". It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

If you're donating, then you are evidently willing to pay without any "buts". So aren't you arguing against your own actions?

Not at all. I donate to Internet Archive, but we're talking here about paying for unobstructed access to but one part of their service: Wayback Machine. Two different things.

Re: An update on Wayback Machine access

#199

Earlier quoted context omitted.

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

> Why folks continue to use them confuses me. I use them. I haven't ever heard mention that the content is edited. Do you have a source?

See

https://arstechnica.com/tech-policy/2026/02/wikipedia-might-...

https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...?

Besides tampering with content, the site was also using visitors to DDOS a blog that mentioned the owner of archive.today.

Re: An update on Wayback Machine access

#200

Earlier quoted context omitted.

Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request. IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

Who do you think the Internet Archive is? They are not Google, they have very low funding, very high expenses and are constantly under legal pressure. The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.

I do donate. But that doesn't mean I have to thank them before every meal or think that the service is perfect.
Post reply on HN