Earlier quoted context omitted.
Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
robots.txt?
An update on Wayback Machine access
291–300 of 374 posts
Re: An update on Wayback Machine access
#292Earlier quoted context omitted.
Once upon a time some people explored backing up the Internet Archive. However, that experiment ended. They mention there were some learnings and they then say: > The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to…
The Internet Archive's torrents are a sick joke. I've yet to find one that actually manages to complete. They always get stuck at 90-something percent but that final blocks always fail verification and get retried, fail, and the process repeats forever. Because they're web seeds they're hitting IA infrastructure and not offloading to a real swarm. So their broken torrents are just screwing themselves.
Re: An update on Wayback Machine access
#293Earlier quoted context omitted.
I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…
Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.
Re: An update on Wayback Machine access
#294Earlier quoted context omitted.
I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…
Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.
It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there.
This is because of their policy of only mirroring what mainstream media outlets are saying.
Re: An update on Wayback Machine access
#295Earlier quoted context omitted.
I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
Are there any decent/reputable residential proxy companies? I’m pretty sure I’ll end up needing one occasionally, for those days when my residential IP has a poor reputation score.
Re: An update on Wayback Machine access
#296I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you…
Re: An update on Wayback Machine access
#297Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.
(I should really start calling rent "tithes" more often)
Re: An update on Wayback Machine access
#298Anyone else feel like making a new internet and starting over
People will just bring their bots over and you'll be back to square one. Bots scraping existed before LLM companies decided to go nutso on the internet.
You could even build this as an overlay on the current internet. DN42 is like this.
Re: An update on Wayback Machine access
#299Re: An update on Wayback Machine access
#300Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.
> Shouldn't the solution be to gate bulk access for automated services for a price? Fine in theory but determined scrapers will use residential proxies in bulk.