Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

221–230 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#221
post #206

Based on this comment: > I definitely get this. The thing that gives me hope is that you only need to poison a very small % of content to damage AI models pretty significantly. It helps combat the mass scraping, because a significant chunk of the data they get will be useless, and its very difficult to filter it by hand It'd be great if the code returned by this project is code that doesn't work. Imagine if all these…

Miasma is just a wrapper around the "Poison Fountain". You can check out the explanation and sample some of their content here: https://rnsaffn.com/poison3/

It's pretty much exactly what you're describing: content that looks correct but is deeply insane.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#222

Earlier quoted context omitted.

The problem I have, is they hammer my site so hard they take it down. The content is for everyone. They can have it. Just don't also take it away from everybody else.

Unintentional denial-of-service attacks from AI scrapers are definitely a problem, I just don't know if "theft" is the right way to classify them. They shouldn't get lumped in with intellectual property concerns, which are a different matter. AI scrapers are a tragedy of the commons problem kind of like Kessler syndrome: a few bad actors can ruin low Earth orbit for everyone via space pollution, which is definitely a…

If I took a photo off your photography blog and used it on my corporate website without your say or input, I don't think it would be unfair to call that stealing.

Doing that on a mass scale with an obfuscation step in between suddenly makes it ok? I'm not convinced.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#223

Earlier quoted context omitted.

Unintentional denial-of-service attacks from AI scrapers are definitely a problem, I just don't know if "theft" is the right way to classify them. They shouldn't get lumped in with intellectual property concerns, which are a different matter. AI scrapers are a tragedy of the commons problem kind of like Kessler syndrome: a few bad actors can ruin low Earth orbit for everyone via space pollution, which is definitely a…

If I took a photo off your photography blog and used it on my corporate website without your say or input, I don't think it would be unfair to call that stealing. Doing that on a mass scale with an obfuscation step in between suddenly makes it ok? I'm not convinced.

[deleted]

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#225

Earlier quoted context omitted.

From my experience, it's the opposite. The more you fuck with them the more they call. It's better not to answer. But I just can't help myself.

I got more calls at first too. I assume they were selling my number to the next scum in line. But then it just stopped.

Its been years though.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#227

I dunno... it feels like the same approach as those people who tell you gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. Also, inserting hidden or misleading links is specifically a no-no for Google Search [0], who have this to say: We detect policy-violating practices both through automated systems and,…

It might work for a very basic bot that doesn't understand how scraping to infinite depth is not very good idea. It won't be effective against anything minimally sophisticated.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#228
post #92

I dunno... it feels like the same approach as those people who tell you gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. Also, inserting hidden or misleading links is specifically a no-no for Google Search [0], who have this to say: We detect policy-violating practices both through automated systems and,…

One would assume legit spiders obey robots.txt.

[flagged]

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#229

Earlier quoted context omitted.

You don't get attribution for your work if it merely feeds into it's training data

That assumes the AI bots are scraping for training data and not simple retrieval/ RAG (which would likely provide attribution)

Oh come on. We know they’re doing both. They are scraping in morally/legally dubious ways as well as doing other things.
Post reply on HN