Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

201–210 of 374 posts

Re: An update on Wayback Machine access

#201

Earlier quoted context omitted.

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

> Why folks continue to use them confuses me. I use them. I haven't ever heard mention that the content is edited. Do you have a source?

See the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment...

They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

Re: An update on Wayback Machine access

#202
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

You might be right but -- why would they have watied until these recent weeks?

Re: An update on Wayback Machine access

#203
post #167

Earlier quoted context omitted.

Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request. IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

What is 'insane' here is the shear level of entitlement displayed here, including lumping a niche, free, volunteer supported service in with billion dollar, for profit corporations and demanding they pander to your inflated expectations. Wild.

It doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.

Re: An update on Wayback Machine access

#206

Earlier quoted context omitted.

Plenty of threads on HN about this, Anubis does not work.

I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

From a few weeks ago: https://news.ycombinator.com/item?id=49500040

Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate.

You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.

Re: An update on Wayback Machine access

#207
post #101

Earlier quoted context omitted.

Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine , it very much is a distinction with a difference.

Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access. So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

I suspect you may be conflating two different things.

Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.

Re: An update on Wayback Machine access

#208

Earlier quoted context omitted.

No - my phone is not connected to work's WiFi. Wonder who the bad actors in my company are...

You should file a support ticket with your manager and the IT security or support desk. Show them the evidence of 429s that are blocking your assigned tasks during working hours. Also include the screenshots and files that you downloaded on your personal device in order to access your work-related materials. Be sure and thank them for adequately configuring the MDM on your personal mobile device so that you could do…

Not sure if you're posting as a joke or sarcasm but accessing archive.org is not relevant to my work.

Re: An update on Wayback Machine access

#209

Why can't they capture OS, Browser, and IP address at the time of error? All that information is available at the point of failure, the user should not need to email it in.

Presumedly they collect that, but the vast majority of blocks are correct and not errors. This info allows them to look up the user’s request to label it as legitimate.

Re: An update on Wayback Machine access

#210

Earlier quoted context omitted.

Who do you think the Internet Archive is? They are not Google, they have very low funding, very high expenses and are constantly under legal pressure. The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.

I do donate. But that doesn't mean I have to thank them before every meal or think that the service is perfect.

What do you expect them to do though? You have to be a reasonable person.
Post reply on HN