Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

81–90 of 374 posts

Re: An update on Wayback Machine access

#81
post #46
post #2

Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon. It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs. I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

Micropayments would solve so many Internet problems. It's not too late to adopt. Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage. The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them.

Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.

Re: An update on Wayback Machine access

#82

Earlier quoted context omitted.

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…

What if they charged money? Is it something you'd pay for?

I'd pay for it, but only if they implemented the changes the community of users have been requesting for years.

Re: An update on Wayback Machine access

#83
post #54

unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?

See also: https://wiki.archiveteam.org/index.php/List_of_websites_excl...

Note that Archive Team is separate from the Internet Archive.

Re: An update on Wayback Machine access

#84
post #74

Earlier quoted context omitted.

Do they offer bulk torrent downloads as an alternative?

I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.

They provide a free cli tool to do this.

Re: An update on Wayback Machine access

#85
post #74

Earlier quoted context omitted.

Do they offer bulk torrent downloads as an alternative?

I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.

isn't that what wget -m does? what is there to pay for?

Re: An update on Wayback Machine access

#87

There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreemen…

Yeah it’s kinda crazy to me that what was once a back alley python script is now accepted as ‘fine, free for all’. The new era of bros really are smth else.

Re: An update on Wayback Machine access

#88
Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?

Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.

Re: An update on Wayback Machine access

#89
post #46

Earlier quoted context omitted.

Micropayments would solve so many Internet problems. It's not too late to adopt. Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage. The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them. Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact…

Micropayments would solve so many problems for the internet. And, cryptocurrencies would solve so many problems for micropayments. But, it's a non-starter because any proposal gets flooded with people popping veins about how crypto can't solve anything.

Re: An update on Wayback Machine access

#90
post #80

Earlier quoted context omitted.

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.

They're different sites, with different goals, run by different people.
Post reply on HN