Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

91–100 of 374 posts

Re: An update on Wayback Machine access

#91

Earlier quoted context omitted.

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…

What if they charged money? Is it something you'd pay for?

I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.

Re: An update on Wayback Machine access

#92
post #80

Earlier quoted context omitted.

ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.

They're different sites, with different goals, run by different people.

That provide the same functional service....

Hence, distinction without a difference.

Re: An update on Wayback Machine access

#93
It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/

I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.

Re: An update on Wayback Machine access

#97

I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.

> Asking users to email them with details of their OS, browser, IP address is just crazy.

It's surely to serve as data to help tell humans apart from bots.

> Changes made by IA shouldn't become my responsibility.

They're a free service. It's ultimately not their responsibility to service you either.

Re: An update on Wayback Machine access

#98
post #93

It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/ I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is li…

I'm a little more skeptical that this is "AI is big so it is worse" issue. Yes, AI is big in scale, but this has been the case for almost every popular free service. They either start:

- charging (news / journalist services)

- gate-keeping (X forcing log-ins)

- enshittifying (lots of ads and degraded service)

The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.

Re: An update on Wayback Machine access

#99
post #92

Earlier quoted context omitted.

They're different sites, with different goals, run by different people.

That provide the same functional service.... Hence, distinction without a difference.

Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.

Re: An update on Wayback Machine access

#100
post #80

Earlier quoted context omitted.

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.

That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
Post reply on HN