Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

181–190 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#181
post #173

Earlier quoted context omitted.

>my comment was about the very human need to be recognized for something created, made, or thought by a person. And I specifically addressed that aspect: >The moral question is more ambiguous, but it's still pretty weak. Most recipes are uncredited, and it's unclear why someone can force everyone to attribute the recipe to them when all they realistically did was tweak the dish a bit. In the example above, I doubt yo…

your ability to not address my argument main point is something to behold. can't tell if you're doing on purpose or not. if humans read my blog posts and then things without credit that would be fine. i like human eyeballs and i like them on my content. that's exactly the purpose of the blog post (_in this particular example_), to get human eyeballs on the content.

>your ability to not address my argument main point is something to behold. can't tell if you're doing on purpose or not.

Or maybe you're just terrible at writing.

>if humans read my blog posts and then things without credit that would be fine.

I'm not sure how I (or anyone) was supposed to come away with this conclusion when you were writing stuff like:

"i'm ok with giving the recipe for free, i just want my name out there"

"the very human need to be recognized for something created"

"they want their name attached and their contribution recognized".

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#182

If you want to ruin someone's web experience based on what kind of thing they are, rather than the content of their character, consider that you might be the baddies.

You’re gonna have to try harder to sneak in the a priori assumption that LLMs have any character beyond which corporation deployed them.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#183
post #181

Earlier quoted context omitted.

your ability to not address my argument main point is something to behold. can't tell if you're doing on purpose or not. if humans read my blog posts and then things without credit that would be fine. i like human eyeballs and i like them on my content. that's exactly the purpose of the blog post (_in this particular example_), to get human eyeballs on the content.

>your ability to not address my argument main point is something to behold. can't tell if you're doing on purpose or not. Or maybe you're just terrible at writing. >if humans read my blog posts and then things without credit that would be fine. I'm not sure how I (or anyone) was supposed to come away with this conclusion when you were writing stuff like: "i'm ok with giving the recipe for free, i just want my name ou…

there is nothing contradictory in what i said, and if you weren't favoring a very literal interpretation of my argument you would agree.

but, in the spirit of critical reading education, what i meant is: human attention good, machine ingestion bad.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#184

Earlier quoted context omitted.

Those are not the only two options. Why are you presenting the latter option as if it were mainstream? It's such a small percentage of use cases that it probably isn't even a rounding error. People who want to disseminate information also want the credit. I'd still like to know why you are presenting this false dichotomy. What reason do you have for presenting a use case that has fractions of a percentage as if it we…

> People who want to disseminate information also want the credit. This is psychological projection.

> This is psychological projection.

You don't know what that means.

In any case, people who want to disseminate information with credit can do so without standing up a blog (any place that allows posting of comments, such as Reddit, HN, etc).

In the context of this discussion, we're talking about site owners; people who put up a blog.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#186

Earlier quoted context omitted.

I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Unless you are Anthropic or OAI it seems like a wild stance to have. It’s good when people are rewarded for works that other people value. If trainers don’t value the work, they shouldn’t train on it. If they do, they should pay for it.

My own view is, I thought we were all agreed that the idea that Microsoft can restrict Wine from even using ideas from Windows, such that people who have read the leaked Windows source cannot contribute to Wine, was a horrible abuse of the legal system that we only went along with under duress? Now when it's our data being used, or more cynically when there's money to be made, suddenly everyone is a copyright maximal…

If you want people to read and learn from each other, you should incentivize people to make content worth reading and learning from. Making LLM training a viable loophole for copyright law means there won’t be incentives to produce such work.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#188

Earlier quoted context omitted.

I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Unless you are Anthropic or OAI it seems like a wild stance to have. It’s good when people are rewarded for works that other people value. If trainers don’t value the work, they shouldn’t train on it. If they do, they should pay for it.

My own view is, I thought we were all agreed that the idea that Microsoft can restrict Wine from even using ideas from Windows, such that people who have read the leaked Windows source cannot contribute to Wine, was a horrible abuse of the legal system that we only went along with under duress? Now when it's our data being used, or more cynically when there's money to be made, suddenly everyone is a copyright maximal…

Re-reading your comment, I think we’re both generally anti-corporate-fuckery. I view the current batch of copyright pearl clutching to be an argument about if VCs are allowed to steal books to make their chatbots worth talking to, and the Wine/MSoft debate about if it should be legal to engage in anticompetitive behavior by restrictive use of copyright. In both of these cases the root of the issue isn’t really the copyright as an abstract- it’s the bludgeoning of the person with less money by use of overwhelming legal costs to have a day in court.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#189
post #73
post #28

> If you have a public website, they are already stealing your work. I have a public website, and web scrapers are stealing my work. I just stole this article, and you are stealing my comment. Thieves, thieves, and nothing but thieves!

I agree theft isn't a good analogy, but there is something similar going on. I put my words out into the world as a form of sharing. I enjoy reading things others write and share freely, so I write so others might enjoy the things I write. But now the things I write and share freely are being used to put money in the bank accounts of the worst people on the planet. They are using my work in a way I don't want it to b…

It sounds like you wanted to believe you were sharing freely while sharing conditionally.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#190
post #155

Earlier quoted context omitted.

Would using an actors face and voice as training data be fair use? What it the model then creates a virtual actor that is very close to the real actor?

>What it the model then creates a virtual actor that is very close to the real actor? "Likeness" is a separate concept from copyrights https://en.wikipedia.org/wiki/Personality_rights

I wish I lived in the alternative timeline where open source folks didn't look a gift horse in the mouth and actually used these tools to copy left the shit out of software to the point where proprietary closed source software has no advantage.

But instead we've got people posting "honey pots" that an LLM will immediately detect and route around.

Post reply on HN