Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
An update on Wayback Machine access
251–260 of 374 posts
Re: An update on Wayback Machine access
#252Re: An update on Wayback Machine access
#253Earlier quoted context omitted.
See the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment... They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.
Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it's fine for gossip and speculation but it shouldn't be in Ars Technica then.
There's also a bunch of previous hackernews discussions about it:
- https://news.ycombinator.com/item?id=47474255
- https://news.ycombinator.com/item?id=46624740
- https://news.ycombinator.com/item?id=47092006
- https://news.ycombinator.com/item?id=46843805
And you obviously have no reason to believe me, but I was following this when it was happening at the start of this year and can confirm that the DDoS script and archive text replacements really did happen.
Re: An update on Wayback Machine access
#254Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.
I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.
Re: An update on Wayback Machine access
#255Re: An update on Wayback Machine access
#256Earlier quoted context omitted.
Examples provided as technical examples, strong feelings are out of scope for this thread.
As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a tempo…
Re: An update on Wayback Machine access
#257Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies? Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.
Re: An update on Wayback Machine access
#258Earlier quoted context omitted.
I thought at least google (and possibly others) provided a way to verify the user agent?
Sorry, not following. I thought the user agent is a string that the caller can set to anything. There is no immediate, reliable way to tell whether the string is correct. Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests ca…
https://developers.google.com/search/blog/2006/09/how-to-ver...
Re: An update on Wayback Machine access
#259Earlier quoted context omitted.
I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.
Is scale what we're discussing though? e.g. a prompt of "fetch and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.
That's the scale argument.