Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

111–120 of 374 posts

Re: An update on Wayback Machine access

#111
post #104

Earlier quoted context omitted.

It's bad form to scrape the scrapers?

Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)

You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason.

The end result is exactly the same.

Re: An update on Wayback Machine access

#112
post #92

Earlier quoted context omitted.

They're different sites, with different goals, run by different people.

That provide the same functional service.... Hence, distinction without a difference.

Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.

Re: An update on Wayback Machine access

#113
post #13

Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

There are definitely factors beyond IP being used. A week ago I found that all requests from Chrome-like browsers got a 429 across more than a half dozen networks and several machines, while Firefox reliably worked. I assume this was an overzealous policy on UAs.

Re: An update on Wayback Machine access

#114

Earlier quoted context omitted.

Examples provided as technical examples, strong feelings are out of scope for this thread.

As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a tempo…

No strong feelings here is what I meant. Certainly, that energy is best directed into aggressive countermeasures and defense in depth of public goods.

https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

Re: An update on Wayback Machine access

#115
post #77

Earlier quoted context omitted.

robots.txt?

Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.

AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.

Re: An update on Wayback Machine access

#116

Are any AI companies using residential proxies to scrape?

Yes. It's impossible to tell which because the split is residential proxies, dataset curators, and AI companies all being separate actors. However I fucking guarantee you it's out there and people are too cowardly to be honest about it so they don't get sued out of existence.

Re: An update on Wayback Machine access

#117

Earlier quoted context omitted.

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

There are definitely factors beyond IP being used. A week ago I found that all requests from Chrome-like browsers got a 429 across more than a half dozen networks and several machines, while Firefox reliably worked. I assume this was an overzealous policy on UAs.

In my work, it's failing on both Firefox and Chrome.

Re: An update on Wayback Machine access

#119

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access. The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gat…

fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future. The future is bleak :\

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side.

It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.

Re: An update on Wayback Machine access

#120
post #65

Earlier quoted context omitted.

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.

I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.
Post reply on HN