Earlier quoted context omitted.
The dream is real, man. If you want open content on the Internet, it's never been a better time. My blog is open to all - machine or man. And it's hosted on my home server next to me. I don't see why anyone would bother trying to distinguish humans from AI. A human hitting your website too much is no different from an AI hitting your website too much. I have a robots.txt that tries to help bots not get stuck in loops…
The only bot that bugs the crap out of me is Anthropic's one. They're the reason I set up a labyrinth using iocaine ( https://iocaine.madhouse-project.org/ ). Their bot was absurdly aggressive, particularly with retries. It's probably trivial in the whole scheme of things, but I love that anthropic spent months making about 10rps against my stupid blog, getting markov chain responses generated from the text of Moby D…
But seriously, Why must someone search even a significant part of the public Internet to develop an AI? Is it believed that missing some text will cripple the AI?
Isn't there some sort of "law of diminishing returns" where, once some percentage of coverage is reached, further scraping is not cost-effective?