Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

71–80 of 374 posts

Re: An update on Wayback Machine access

#71
post #46
post #2

Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon. It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs. I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

Micropayments would solve so many Internet problems. It's not too late to adopt. Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage. The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

Micropayments are blocked by government money laundering regulations increasing the costs significantly to make them untenable.

Re: An update on Wayback Machine access

#72

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access. The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gat…

fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future.

The future is bleak :\

Re: An update on Wayback Machine access

#73
post #65

Earlier quoted context omitted.

As it should. Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

robots.txt?

Re: An update on Wayback Machine access

#74
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

Do they offer bulk torrent downloads as an alternative?

I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.

Re: An update on Wayback Machine access

#75
There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?

Some approaches that I think are promising:

- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).

- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.

- what else?

[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...

[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...

Re: An update on Wayback Machine access

#76
I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.

Re: An update on Wayback Machine access

#77
post #65

Earlier quoted context omitted.

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

robots.txt?

Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.

Re: An update on Wayback Machine access

#79
post #40

[flagged]

Yes, that's one acceptable alternative, and another commonly accepted alternative is API's. Although, I'm not sure why you included the asterisk.

Unkeen on the apostrophe that’s why! Gotta keep HN proper and correct guize

Re: An update on Wayback Machine access

#80

Earlier quoted context omitted.

Every single paid article linked on HN has the way back machine link as the very first comment.

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
Post reply on HN