Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

361–370 of 374 posts

Re: An update on Wayback Machine access

#361

Earlier quoted context omitted.

Do you have a library card?

I can walk into a library, pull a book from a shelf, sit down, and read it front to back. No library card needed. Only need one to take a book home. Completely different scenario. Edit:clarification

Only because there's no way for bots to exploit the system since it's in the physical world.

If thousands of robots suddenly showed up at your local library and started hogging all the books so nobody else could use the library, you can be sure a library card would be required to even enter.

Re: An update on Wayback Machine access

#362
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…

>I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice

They didn't use to, this has become a thing over the past few years as MSM outlets have completely given up on journalistic standards, including editorial ones.

Re: An update on Wayback Machine access

#363

Earlier quoted context omitted.

People will just bring their bots over and you'll be back to square one. Bots scraping existed before LLM companies decided to go nutso on the internet.

You can try a higher level of identity verification on the new internet. It shouldn't be fully ID verified, but more like how it used to be - users on a network were anonymous to other networks, but you could email the admin of a network to track down bad behavior with their cooperation if they agreed it was bad. You could even build this as an overlay on the current internet. DN42 is like this.

So then identity theft and fake identities will sky-rocket.

Re: An update on Wayback Machine access

#364
post #361

Earlier quoted context omitted.

I can walk into a library, pull a book from a shelf, sit down, and read it front to back. No library card needed. Only need one to take a book home. Completely different scenario. Edit:clarification

Only because there's no way for bots to exploit the system since it's in the physical world. If thousands of robots suddenly showed up at your local library and started hogging all the books so nobody else could use the library, you can be sure a library card would be required to even enter.

You clearly do not live in an area with a large homeless population. Many libraries are effectively under-resourced homeless shelters. Yet, no requirement to have a library card to get in. Information is still open to the public.

Re: An update on Wayback Machine access

#365

Earlier quoted context omitted.

Generally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible. Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible…

we’ve been using Datadome at work for this because it’s nice to get someone else to think about the constant bot cat and mouse, and we all benefit from rules and fixes created from other client data. Not an ad- that service is eye watering expensive but I think it makes sense as something to offload.

How much is eye watering expensive roughly?

Re: An update on Wayback Machine access

#366

Earlier quoted context omitted.

As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a tempo…

I say put your data up in torrents, host a few KB of plain HTML linking to them, and let decentralisation do the rest.

The data IS available. You can download it all from several sources in one big dump. And yet we're still scraped.

Re: An update on Wayback Machine access

#367

Earlier quoted context omitted.

I say put your data up in torrents, host a few KB of plain HTML linking to them, and let decentralisation do the rest.

The data IS available. You can download it all from several sources in one big dump. And yet we're still scraped.

Further evidence that they're not actually going after your data, but just DDoS'ing.

Re: An update on Wayback Machine access

#368

I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.

Plenty of threads on HN about this, Anubis does not work.

I am thankful for Anubis. It's a litmus test for the arrogance of the people behind a website.

Re: An update on Wayback Machine access

#369

Earlier quoted context omitted.

I literally had never heard of this before. I don't check HN every single day. It's extremely reasonable to ask for a link, very easy to include one when making a claim, and attacking someone for asking for evidence is extremely anti-intellectual independent of the level of effort required. (n.b. that doesn't excuse the hostile way that they asked for proof - "Let's all just believe this baseless assertion shall we")

[dead]

They were pointing out the lack of evidence on my part. I agree it was rude but I don't think it's helpful to start calling it misinformation with no evidence. They had a valid point that not everybody Just Knows already, hence why I did reply with a link. I don't think it's constructive to jab much more than I did in that reply.

Re: An update on Wayback Machine access

#370

Earlier quoted context omitted.

What do you expect them to do though? You have to be a reasonable person.

Data analysis would be a good start, better blocking heuristics, an off-the-shelf solution used by other organisations that don't have this problem, etc.

[dead]
Post reply on HN