Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

51–60 of 374 posts

Re: An update on Wayback Machine access

#51
post #32

Earlier quoted context omitted.

Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy. On the other hand, the Internet Archive is a non-profit offering a free public resource.

Examples provided as technical examples, strong feelings are out of scope for this thread.

As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.

The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

[0] https://people.kernel.org/monsieuricon/creepy-crawlies

Re: An update on Wayback Machine access

#52
post #45
post #35

AI companies should pay billions to wayback machine for access

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?

Re: An update on Wayback Machine access

#53

[flagged]

No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.

There are people commenting parallel to you saying they are upset about AI companies scraping the web.

Re: An update on Wayback Machine access

#55
post #45
post #35

AI companies should pay billions to wayback machine for access

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.

Re: An update on Wayback Machine access

#56
post #49
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

Just yesterday from my one of my sessions with Sol: > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

As it should.

Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

Re: An update on Wayback Machine access

#57
post #52
post #45

Earlier quoted context omitted.

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?

[deleted]

Re: An update on Wayback Machine access

#58
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

Re: An update on Wayback Machine access

#59
post #52
post #45

Earlier quoted context omitted.

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?

[flagged]
Post reply on HN