Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

111–120 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#111

Earlier quoted context omitted.

search crawlers used to bring people TO your site llm boots are used to keep people OUT of your site, because knowledge is indexed and distributed by corporations.

So if your site is dependent on ads, and since the only way for people to see those ads is coming to your site, then yes, you lose. If your site exists to share information, then the information gets disseminated, whether via LLM or some browser, it doesn't make a difference to me

You don't get attribution for your work if it merely feeds into it's training data

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#112

Earlier quoted context omitted.

So if your site is dependent on ads, and since the only way for people to see those ads is coming to your site, then yes, you lose. If your site exists to share information, then the information gets disseminated, whether via LLM or some browser, it doesn't make a difference to me

You don't get attribution for your work if it merely feeds into it's training data

That assumes the AI bots are scraping for training data and not simple retrieval/ RAG (which would likely provide attribution)

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#113

Earlier quoted context omitted.

search crawlers used to bring people TO your site llm boots are used to keep people OUT of your site, because knowledge is indexed and distributed by corporations.

So if your site is dependent on ads, and since the only way for people to see those ads is coming to your site, then yes, you lose. If your site exists to share information, then the information gets disseminated, whether via LLM or some browser, it doesn't make a difference to me

Those are not the only two options.

Why are you presenting the latter option as if it were mainstream? It's such a small percentage of use cases that it probably isn't even a rounding error.

People who want to disseminate information also want the credit.

I'd still like to know why you are presenting this false dichotomy. What reason do you have for presenting a use case that has fractions of a percentage as if it were a standard use case? What is your motivation behind this?

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#114

Earlier quoted context omitted.

What "content of character" do you ascribe to a web scraper?

You don't, that's why it's unethical to block them. If you keep getting harrassed by people wearing black hoodies, would it be ethical to start taking countermeasures against all people who wear black hoodies?

If they are coming to my door to harass me, then yes, it makes sense to take countermeasures against all black-hoodie wearers when I see them at the door.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#115
post #90

Earlier quoted context omitted.

> If you put stuff out in public for anyone to use, then find out it's used in a way you don't like Nope. Copyright is a thing, licenses are a thing. Both are completely ignored by LLM companies, which was already proven in court, and for which they already had to pay billions in fines. Just because something is publicly accessible, that does not mean everybody is entitled to abuse it for everything they see fit.

>Nope. Copyright is a thing, licenses are a thing. Both are completely ignored by LLM companies, which was already proven in court, ...the same courts that ruled that AI training is probably fair use? Fair use trumps whatever restrictions author puts on their "licenses". If you're an author and it turned out that your book was pirated by AI companies then fair enough, but "I put my words out into the world as a form…

I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Unless you are Anthropic or OAI it seems like a wild stance to have. It’s good when people are rewarded for works that other people value. If trainers don’t value the work, they shouldn’t train on it. If they do, they should pay for it.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#116
post #28

> If you have a public website, they are already stealing your work. I have a public website, and web scrapers are stealing my work. I just stole this article, and you are stealing my comment. Thieves, thieves, and nothing but thieves!

The problem I have, is they hammer my site so hard they take it down. The content is for everyone. They can have it. Just don't also take it away from everybody else.

Unintentional denial-of-service attacks from AI scrapers are definitely a problem, I just don't know if "theft" is the right way to classify them. They shouldn't get lumped in with intellectual property concerns, which are a different matter. AI scrapers are a tragedy of the commons problem kind of like Kessler syndrome: a few bad actors can ruin low Earth orbit for everyone via space pollution, which is definitely a problem, but saying that they "stole" LEO from humanity doesn't feel like the right terminology. Maybe the problem with AI scrapers could be better described as "bandwidth pollution" or "network overfishing" or something.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#117

Way back in the day I had a software product, with a basic system to prevent unauthorised sharing, since there was a small charge for it. Every time I released an update, and new crack would appear. For the next six months I worked on improving the anti-copying code until I stumbled across an article by a coder in the same boat as me. He realised he was now playing a game with some other coders where he make the copy…

> the cracker would then have fun cracking it.

I wonder if you could've won by making the cracking boring. No new techniques, bare minimum changes to require compiling a new crack, and just enough to make it difficult to automate. I.e. turn the cracking into a job.

But in reality, there are other community-driven motivations to put out cracks.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#118

I dunno... it feels like the same approach as those people who tell you gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. Also, inserting hidden or misleading links is specifically a no-no for Google Search [0], who have this to say: We detect policy-violating practices both through automated systems and,…

yes it work.

phone scammers have a very high personel cost, hence why some resort for human traffic.

if everyone picked up the phone and wasted a few seconds, it would be enough to make their whole enterprise worthless. but since most people who would not fail shutdown right away, they have the best ROI of any industry. they don't even pay the call for first seconds.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#119

Earlier quoted context omitted.

The problem I have, is they hammer my site so hard they take it down. The content is for everyone. They can have it. Just don't also take it away from everybody else.

Unintentional denial-of-service attacks from AI scrapers are definitely a problem, I just don't know if "theft" is the right way to classify them. They shouldn't get lumped in with intellectual property concerns, which are a different matter. AI scrapers are a tragedy of the commons problem kind of like Kessler syndrome: a few bad actors can ruin low Earth orbit for everyone via space pollution, which is definitely a…

you're totally right about not being theft, but we have a term. you used it yourself, "distributed denial of service". that's all it is. these crawlers should be kicked off the internet for abuse. people should contact the isp of origin.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#120
post #90

Earlier quoted context omitted.

>Nope. Copyright is a thing, licenses are a thing. Both are completely ignored by LLM companies, which was already proven in court, ...the same courts that ruled that AI training is probably fair use? Fair use trumps whatever restrictions author puts on their "licenses". If you're an author and it turned out that your book was pirated by AI companies then fair enough, but "I put my words out into the world as a form…

I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Unless you are Anthropic or OAI it seems like a wild stance to have. It’s good when people are rewarded for works that other people value. If trainers don’t value the work, they shouldn’t train on it. If they do, they should pay for it.

>I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training.

Fair use is part of "copyright and licensing laws".

Post reply on HN