Earlier quoted context omitted.
Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.
AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.
An update on Wayback Machine access
181–190 of 374 posts
Re: An update on Wayback Machine access
#182> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…
Re: An update on Wayback Machine access
#183Earlier quoted context omitted.
The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
I use them. I haven't ever heard mention that the content is edited. Do you have a source?
Re: An update on Wayback Machine access
#184Earlier quoted context omitted.
fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future. The future is bleak :\
I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…
Re: An update on Wayback Machine access
#185Earlier quoted context omitted.
It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.
I'd have a lot less of a problem with AI if everything that went into their training was public domain and made easily available to anyone for any use. It'd feel less like AI companies were just stealing the work of others and charging for it.
* If you train AI on it, you have to afford public access to it.
* Nobody can exact violence against anybody else in response to that person providing public access to any data anymore (ie, all bytestrings are public domain).
That's the world I'd like to try in the coming years.
Re: An update on Wayback Machine access
#186Earlier quoted context omitted.
I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.
Is scale what we're discussing though? e.g. a prompt of "fetch and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.
The problem is that it is hard to distinguish your one off (which seems perfectly fine) from the tidal wave of bad actors.
Re: An update on Wayback Machine access
#187Earlier quoted context omitted.
I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.
I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.
Re: An update on Wayback Machine access
#188Earlier quoted context omitted.
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
It's an open secret that you can often circumvent paywalls by searching Wayback.
Re: An update on Wayback Machine access
#189Earlier quoted context omitted.
Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason . The end result is exactly the same.
Re: An update on Wayback Machine access
#190Earlier quoted context omitted.
Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine , it very much is a distinction with a difference.
Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access. So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.
The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".