Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

151–160 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#151
post #120

Earlier quoted context omitted.

I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Unless you are Anthropic or OAI it seems like a wild stance to have. It’s good when people are rewarded for works that other people value. If trainers don’t value the work, they shouldn’t train on it. If they do, they should pay for it.

>I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Fair use is part of "copyright and licensing laws".

Would using an actors face and voice as training data be fair use?

What it the model then creates a virtual actor that is very close to the real actor?

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#152
post #129

Earlier quoted context omitted.

>Is it ok to get the recipe, remove my name, and write in LLM-Codex as the creator? again, i'm ok with giving the recipe for free, i just want my name out there. From a legal perspective, it's a pretty clear "no". The instructions in recipes aren't copyrightable. The moral question is more ambiguous, but it's still pretty weak. Most recipes are uncredited, and it's unclear why someone can force everyone to attribute…

i'm curious, do you honestly think the argument was about recipes and cookies? maybe it was an analogy? looking back up the comment tree, it does seem to be an analogy, not a discussion about ACTUAL cookies and ACTUAL recipes.

>maybe it was an analogy?

In that case it's a terrible analogy because if you can't get people to agree on the cookies case, what hope do you have to extend it to the case you're trying to apply the analogy to? It's like saying "You wouldn't pirate a movie, why would you pirate a blog post", because most people would pirate movies.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#154

Earlier quoted context omitted.

If someone hands out cookies in the supermarket, are you allowed to grab everything and leave?

Odd thing about cookies… they disappear after one serving. Websites are an endless stream of cookies. The analogy doesn’t hold.

[deleted]

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#155
post #120

Earlier quoted context omitted.

>I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Fair use is part of "copyright and licensing laws".

Would using an actors face and voice as training data be fair use? What it the model then creates a virtual actor that is very close to the real actor?

>What it the model then creates a virtual actor that is very close to the real actor?

"Likeness" is a separate concept from copyrights

https://en.wikipedia.org/wiki/Personality_rights

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#156

Earlier quoted context omitted.

style="display: none;" aria-hidden="true" tabindex="1" many scraper already know not to follow these, as it's how site used to "cheat" pagerank serving keyword soups

You dont have to use this. You can have it visible bit hide it for humans with other easy tricks.

Scrapers can work around those other easy tricks too.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#157
post #145

Earlier quoted context omitted.

Firstly, since this argument is about semantic pedantry anyways, it's just denial-of-service, not distributed denial-of-service. AI scraper requests come from centralized servers, not a botnet. Secondly, denial-of-service implies intentionality and malice that I don't think is present from AI scrapers. They cause huge problems, but only as a negligent byproduct of other goals. I think that the tragedy of the commons…

The first is incorrect, these scrapers are usually distributed across many IPs, in my experience. I usually refer to them as "disturbed, non-identifying crawlers (DNCs)" when I want to be maximally explicit. (The worst I've seen is some crawler/botnet making exactly one request per IP -_-)

I think the second is incorrect too. DDoS is a DDoS no matter what the intent is.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#159
post #152

Earlier quoted context omitted.

i'm curious, do you honestly think the argument was about recipes and cookies? maybe it was an analogy? looking back up the comment tree, it does seem to be an analogy, not a discussion about ACTUAL cookies and ACTUAL recipes.

>maybe it was an analogy? In that case it's a terrible analogy because if you can't get people to agree on the cookies case, what hope do you have to extend it to the case you're trying to apply the analogy to? It's like saying "You wouldn't pirate a movie, why would you pirate a blog post", because most people would pirate movies.

oh man.

my comment was about the very human need to be recognized for something created, made, or thought by a person. People are ok with writing blog posts, they're ok with writing software, and they're ok with give it all for free, but they want their name attached and their contribution recognized.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#160

Earlier quoted context omitted.

If someone hands out cookies in the supermarket, are you allowed to grab everything and leave?

Odd thing about cookies… they disappear after one serving. Websites are an endless stream of cookies. The analogy doesn’t hold.

Fine.

Me and my 9 friends stand around the cookie-serving person blocking everyone else.

It's taking all the cookies over a period of time.

The analogy was good.

Post reply on HN