Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

101–110 of 374 posts

Re: An update on Wayback Machine access

#101
post #92

Earlier quoted context omitted.

That provide the same functional service.... Hence, distinction without a difference.

Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine , it very much is a distinction with a difference.

Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access.

So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

Re: An update on Wayback Machine access

#102
post #92

Earlier quoted context omitted.

They're different sites, with different goals, run by different people.

That provide the same functional service.... Hence, distinction without a difference.

They are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc.

I think it's a distinction worth making.

Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.

Re: An update on Wayback Machine access

#103

Earlier quoted context omitted.

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…

What if they charged money? Is it something you'd pay for?

I wonder if there would be concern on their part about appearing to be a company that was basically offering paywall circumvention as a product.

Re: An update on Wayback Machine access

#104
post #80

Earlier quoted context omitted.

ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.

That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.

It's bad form to scrape the scrapers?

Re: An update on Wayback Machine access

#106
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

Do they offer bulk torrent downloads as an alternative?

Once upon a time some people explored backing up the Internet Archive.

However, that experiment ended. They mention there were some learnings and they then say:

> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.

https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK

I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.

I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.

Re: An update on Wayback Machine access

#107
post #13

Wow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

No - my phone is not connected to work's WiFi.

Wonder who the bad actors in my company are...

Re: An update on Wayback Machine access

#108
post #104

Earlier quoted context omitted.

That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.

It's bad form to scrape the scrapers?

Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)

Re: An update on Wayback Machine access

#109
post #54

unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?

They get marked as inaccessible, but still exist in IA data

is it possible to access somehow? it seems the site got excluded because of robots.txt set by some domain squatter, not manually.

Re: An update on Wayback Machine access

#110
post #65

Earlier quoted context omitted.

As it should. Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.
Post reply on HN