Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

291–300 of 374 posts

Re: An update on Wayback Machine access

#291
post #65

Earlier quoted context omitted.

Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.

robots.txt?

robots.txt is a shitshow just like user agents. It's been twisted so many ways it doesn't reliably signal actual intent any more.

Re: An update on Wayback Machine access

#292

Earlier quoted context omitted.

Once upon a time some people explored backing up the Internet Archive. However, that experiment ended. They mention there were some learnings and they then say: > The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to…

The Internet Archive's torrents are a sick joke. I've yet to find one that actually manages to complete. They always get stuck at 90-something percent but that final blocks always fail verification and get retried, fail, and the process repeats forever. Because they're web seeds they're hitting IA infrastructure and not offloading to a real swarm. So their broken torrents are just screwing themselves.

The torrent generator races with the uploader and runs on a half-finished upload.

Re: An update on Wayback Machine access

#293
post #287

Earlier quoted context omitted.

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…

Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.

Strangers can't really edit Wikipedia any more, especially if their edit is suspicious. It's a closed system despite the appearance. An anti-vandal bot or human will quickly revert your edit.

Re: An update on Wayback Machine access

#294

Earlier quoted context omitted.

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…

Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.

This became obvious to me when from 2023-2025 they refused to call it anything other than "Israel-Hamas war"

It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there.

This is because of their policy of only mirroring what mainstream media outlets are saying.

Re: An update on Wayback Machine access

#295
post #225

Earlier quoted context omitted.

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

Are there any decent/reputable residential proxy companies? I’m pretty sure I’ll end up needing one occasionally, for those days when my residential IP has a poor reputation score.

No, they are all grey-market. Some have more professional-looking websites, but they're all using the same proxies. This should not prevent you from using them.

Re: An update on Wayback Machine access

#296
post #17

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you…

I wonder if your browser is prefetching every link you move the mouse over.

Re: An update on Wayback Machine access

#297

Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.

by the enclosure "movement"? i.e. greedy powerful people walling off everything they could and declaring it was theirs and you'd have to pay a tithe to use it?

(I should really start calling rent "tithes" more often)

Re: An update on Wayback Machine access

#298

Anyone else feel like making a new internet and starting over

People will just bring their bots over and you'll be back to square one. Bots scraping existed before LLM companies decided to go nutso on the internet.

You can try a higher level of identity verification on the new internet. It shouldn't be fully ID verified, but more like how it used to be - users on a network were anonymous to other networks, but you could email the admin of a network to track down bad behavior with their cooperation if they agreed it was bad.

You could even build this as an overlay on the current internet. DN42 is like this.

Re: An update on Wayback Machine access

#299
post #95

I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.

Gatekeeping information is not the solution

Do you have a library card?

Re: An update on Wayback Machine access

#300

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.

> Shouldn't the solution be to gate bulk access for automated services for a price? Fine in theory but determined scrapers will use residential proxies in bulk.

Make a user download and hash 100GB of junk data before being allowed in. Residential proxies cost a lot per GB.
Post reply on HN