Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

251–260 of 374 posts

Re: An update on Wayback Machine access

#251
post #13

Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

My buddy said he could not access it even from a residential IP, it was blacklisted for some reason.

Re: An update on Wayback Machine access

#252

Earlier quoted context omitted.

One site was doing that which archive in their name but wasn't apart of archive.org

archive.is or archive.today? Why being shy in the era of stealing AI?

archive.today uses clients to perform DDoS, I would not recommend using their site.

Re: An update on Wayback Machine access

#253
post #211
post #201

Earlier quoted context omitted.

See the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment... They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it's fine for gossip and speculation but it shouldn't be in Ars Technica then.

There is a megalodon.jp archive of an archive.today page that shows the edits: https://megalodon.jp/2026-0219-1654-07/https://archive.ph:44...

There's also a bunch of previous hackernews discussions about it:

- https://news.ycombinator.com/item?id=47474255

- https://news.ycombinator.com/item?id=46624740

- https://news.ycombinator.com/item?id=47092006

- https://news.ycombinator.com/item?id=46843805

And you obviously have no reason to believe me, but I was following this when it was happening at the start of this year and can confirm that the DDoS script and archive text replacements really did happen.

Re: An update on Wayback Machine access

#254
The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic

Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.

I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.

Re: An update on Wayback Machine access

#256

Earlier quoted context omitted.

Examples provided as technical examples, strong feelings are out of scope for this thread.

As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a tempo…

I say put your data up in torrents, host a few KB of plain HTML linking to them, and let decentralisation do the rest.

Re: An update on Wayback Machine access

#257

Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies? Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.

It's not. There is no evidence, just propaganda.

Re: An update on Wayback Machine access

#258

Earlier quoted context omitted.

I thought at least google (and possibly others) provided a way to verify the user agent?

Sorry, not following. I thought the user agent is a string that the caller can set to anything. There is no immediate, reliable way to tell whether the string is correct. Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests ca…

Google's scraper bot at least used to be behind IPs that you could identify via reverse-then-forward DNS. Not sure if that is still up to date, though.

https://developers.google.com/search/blog/2006/09/how-to-ver...

Re: An update on Wayback Machine access

#259

Earlier quoted context omitted.

I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.

Is scale what we're discussing though? e.g. a prompt of "fetch and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

Sure, but all the time I'll ask Claude a question, and then I'll see it fetch 5-10 different URLs to come up with answer. I certainly would not be fetching those URLs at that rate if I were doing it myself. I would probably be visiting those pages, one by one, over the span of 10-20 minutes.

That's the scale argument.

Re: An update on Wayback Machine access

#260

Earlier quoted context omitted.

archive.today uses clients to perform DDoS, I would not recommend using their site.

[flagged]

https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...> for those without a search engine.
Post reply on HN