Earlier quoted context omitted.
If someone hands out cookies in the supermarket, are you allowed to grab everything and leave?
This is a dishonest analogy. In your example, there is only a limited amount of cookies available. While there is no practical limit on the amount of time a certain digital media can be viewed. You are allowed to take one cookie. But you are allowed to view a public website multiple times if you so want.
Miasma: A tool to trap AI web scrapers in an endless poison pit
61–70 of 276 posts
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#62The irony of machine-generated slop to fight machine-generated slop would be funny, if it weren't for the implications. How long before people start sharing ai-spam lists, both pro-ai and anti-ai?
Just like with email, at some point these share-lists will be adopted by the big corporates, and just like with email will make life hard for the small players.
Once a website appears on one of these lists, legitimately or otherwise, what'll be the reputational damage hurting appearance in search indexes? There have already been examples of Google delisting or dropping websites in search results.
Will there be a process to appeal these blacklists? Based on how things work with email, I doubt this will be a meaningful process. It's essentially an arms race, with the little folks getting crushed by juggernauts on all sides.
This project's selective protection of the major players reinforces that effect; from the README:
" Be sure to protect friendly bots and search engines from Miasma in your robots.txt!
User-agent: Googlebot User-agent: Bingbot User-agent: DuckDuckBot User-agent: Slurp User-agent: SomeOtherNiceBot Disallow: /bots Allow: / "
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#63If you want to ruin someone's web experience based on what kind of thing they are, rather than the content of their character, consider that you might be the baddies.
It's not all that productive, it's an act of desperation. If you can't stop the enemy, at least you can make their action more costly.
One positive outcome I could see it AI companies becoming more critical of their training data.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#64Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#65Earlier quoted context omitted.
If someone hands out cookies in the supermarket, are you allowed to grab everything and leave?
This is a dishonest analogy. In your example, there is only a limited amount of cookies available. While there is no practical limit on the amount of time a certain digital media can be viewed. You are allowed to take one cookie. But you are allowed to view a public website multiple times if you so want.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#66Is there any evidence or hints that these actually work? It seems pretty reasonable that any scraper would already have mitigations for things like this as a function of just being on the internet.
It might work against people just use their Mini Mac with OpenClaw to summarize news every morning, but it certainly won't work against Google. More centralized web ftw.
Good enough for me.
> More centralized web ftw.
This ain't got anything to do with "centralized web," this kind of epistemological vandalism can't be shunned enough.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#67> If you have a public website, they are already stealing your work. I have a public website, and web scrapers are stealing my work. I just stole this article, and you are stealing my comment. Thieves, thieves, and nothing but thieves!
If someone hands out cookies in the supermarket, are you allowed to grab everything and leave?
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#68My asthmar I'm assuming this is a reference to Lord of the flies
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#69Isn't it the case that AI models learn better and are more performant with carefully curated material, so companies do actually filter for quality input?
Isn't it also the case that the use of RLHF and other refinement techniques essentially 'cures' the models of bad input?
Isn't it also, potentially, the case that the ai-scrapers are mostly looking for content based on user queries, rather than as training data?
If the answers to the questions lean a particular way (yes to most), then isn't the solution rate-limiting incoming web-queries rather than (presumed) well-poisoning?
Is this a solution in search of a problem?
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#70Is there any evidence or hints that these actually work? It seems pretty reasonable that any scraper would already have mitigations for things like this as a function of just being on the internet.
I'm completely uncertain that the unsophisticated garbage I generated makes any difference, much less "poisons" the LLMs. A fellow can dream, can't he?