Is there any evidence or hints that these actually work? It seems pretty reasonable that any scraper would already have mitigations for things like this as a function of just being on the internet.
About two years ago, I made up reference to a nonexistent python library and put code "using" it in just 5 GitHub repos. Several months later the free ChatGPT picked it up. So IMO it works.
Miasma: A tool to trap AI web scrapers in an endless poison pit
101–110 of 276 posts
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#102Why do this though? It's like if someone was trying to "trap" search crawlers back in the early 2000s. Seems counterproductive
Because of bots that don't respect ROBOTS.txt . If you want an AI bot to crawl your website while you pay for that bandwidth then you wont use the tool.
Like, what if you actually post something that gains traction, is it going to bankrupt you or something?
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#103> If you have a public website, they are already stealing your work. I have a public website, and web scrapers are stealing my work. I just stole this article, and you are stealing my comment. Thieves, thieves, and nothing but thieves!
I agree theft isn't a good analogy, but there is something similar going on. I put my words out into the world as a form of sharing. I enjoy reading things others write and share freely, so I write so others might enjoy the things I write. But now the things I write and share freely are being used to put money in the bank accounts of the worst people on the planet. They are using my work in a way I don't want it to b…
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#104If you want to ruin someone's web experience based on what kind of thing they are, rather than the content of their character, consider that you might be the baddies.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#105Earlier quoted context omitted.
> This is ultimately just going to give them training material for how to avoid this crap. > The arms race just took another step, and if you're spending money creating or hosting this kind of content, it's not going to make up for the money you're losing by your other content getting scraped. So we should all just do nothing and accept the inevitable?
> So we should all just do nothing and accept the inevitable? I daresay rate-limiting will result in better outcomes than well-poisoning with hidden links that are against the policies of search engines. Lots of potential for collateral damage, including your own websites' reputations and search visibility, with the well-poisoning approach.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#106Why do this though? It's like if someone was trying to "trap" search crawlers back in the early 2000s. Seems counterproductive
Web crawlers didn’t routinely take down public resources or use the scraped info to generate facsimiles that people are still having ethical debates over. Its presence didn’t even register and it was indexing that helped them. It isn’t remotely the same thing. https://www.libraryjournal.com/story/ai-bots-swarm-library-c...
And search crawlers/results have been producing snippets that prevent users from clicking to the source for well over a decade.
Edit: it loaded. I don't see how the problem isn't simply solved by an off the shelf solution like cloud flare. In the real world, you wouldn't open up a space/location if you couldn't handle the throughput. Why should online spaces/locations get special treatment?
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#107Isn't this a trope at this point? That AI companies are indiscriminately training on random websites? Isn't it the case that AI models learn better and are more performant with carefully curated material, so companies do actually filter for quality input? Isn't it also the case that the use of RLHF and other refinement techniques essentially 'cures' the models of bad input? Isn't it also, potentially, the case that t…
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#108Earlier quoted context omitted.
If someone hands out cookies in the supermarket, are you allowed to grab everything and leave?
Odd thing about cookies… they disappear after one serving. Websites are an endless stream of cookies. The analogy doesn’t hold.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#109Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#110If you want to ruin someone's web experience based on what kind of thing they are, rather than the content of their character, consider that you might be the baddies.
What "content of character" do you ascribe to a web scraper?
If you keep getting harrassed by people wearing black hoodies, would it be ethical to start taking countermeasures against all people who wear black hoodies?