Live data from Hacker News

Fear of AI just killed a useful tool

techdirt.com

221–230 of 309 posts

Re: Fear of AI just killed a useful tool

#221
post #122
post #114

Earlier quoted context omitted.

How was the tool in question, the tool the book authors were outraged about, "stealing their IP"? It's also unclear how it would make any money at all -- indeed, the tool's author even said it made no money.

I believe that the authors are not fighting against that tool in particular. They are fighting about generative AIs being trained with their material without their consent. Which is totally legit to me. Maybe this tool is just a collateral damage of a much bigger debate, I don't know. The fact is that in the bigger debate, it does feel like engineers don't seem to care much about the artists. Why would the artists ca…

[deleted]

Re: Fear of AI just killed a useful tool

#223
post #140
post #93

Earlier quoted context omitted.

> The article itself is clueless… it doesn’t engage authors’ concerns at all, and just portrays authors as stoopid AI-fearful luddites. Going off of some of the tweets about this that initially whipped up the outrage about this…it’s not like they were making a nuanced case about their concerns, they were basically just stomping their feet and shouting.

That's twitter generally. If your engagement only reaches the level of twitter, you aren't really engaging at all.

So as long as that's all the engagement there is, we're free to ignore it and carry on, correct?

Re: Fear of AI just killed a useful tool

#226

Earlier quoted context omitted.

While I would agree in theory that a project like this would be best with opt-in, in reality that would just not work. Publishers would never opt-in to it, if they even respond to your requests at all.

Then don't do it? Or, if you do it, do it privately and don't share it on the internet? I'm not sure why this is a difficult idea; if asking for something and getting permission to do it is so difficult that 'would just not work. Publishers would never opt-in to it' ...then, it seems really obvious that even if you want to do it, technically can do it and you could maybe make a legal argument to doing it doesn't viol…

Because copyright in fact is not that strict (Google Books does far more) and you don’t need to respect someone’s boundaries when they don’t have a legal right to those boundaries. Why should we sympathize with people who want far stricter control over the cultural commons?

Re: Fear of AI just killed a useful tool

#227

I believe that if you are using my book as data/model/whatever for something that makes money, then I want a piece of the action. If it's just a paragraph, fine go ahead. But the whole book? Give me money.

So you want money from people who read books and writes reviews? I don't think that's reasonable.

Re: Fear of AI just killed a useful tool

#228
post #127

Earlier quoted context omitted.

Sorry I did not understand that :-). My point was that it is different: when humans read a book, they don't train a machine learning model. They can't read as many books as a machine, at the same speed, and they can't remember nearly as much as what a machine can. Humans and computers are fundamentally different, and it matters. You can't conclude that because it works for one, it will fork for the other.

> Sorry I did not understand that :-) You seemed to be saying that the differences I listed (quicker and more specific feedback) were the only differences. Those are both positive. I was saying that some people may think there are negative differences as well.

Right. Yeah I did not express myself clearly, sorry :). You were saying "how is it different other than X and Y?", and I wanted to say that X and Y are already enough for me to consider them different.

I am actually on the side that LLMs are a big problem for copyright, and I don't want my code and blog posts to be used in their training dataset without my consent. To me, at this scale, it's not fair use. IMO it's a bit like if Facebook said that it is fair use to leverage metadata about their users, because "someone who sees you in a public space talking to a friend knows that you are talking with that person, and it is the same for Facebook on social media". My problem is not that Facebook knows that I sent a message to a friend now, but rather that they know who writes to whom and when, at scale.

Similarly my problem is not that somebody could read my blog post, learn from it, and write another blog post. My problem is that LLMs automatically train on all written material they want on the Internet, at scale, and without acknowledging that all that material has a lot of value (and is copyrighted).

I think fair use should somehow consider the scale.

Re: Fear of AI just killed a useful tool

#229
The tool looks like it was useful to a certain kind of person. If it would actually make money, I would gladly (and probably easily) replicate it because I don't really care that much if Internet randos hate me. But I don't have a good idea for how to keep the return / effort ratio high on this.

I don't need consent for a lot of this, and I probably wouldn't bother. If I made a "List of books with terrible sentences" I wouldn't ask for opt-in or even bother contacting the authors. I will just make the list and quote the sentence.

The law and public opinion is on my side, though I only need the former.

Re: Fear of AI just killed a useful tool

#230
post #149
post #134

Earlier quoted context omitted.

If you want to be a pitchfork mob against generative AI at least understand whether AI is generative or not? Seems like a reasonably low bar. This was non-generative AI, it didn't produce content it output metrics and labelled some existing content.

What makes you think that I don't understand whether AI is generative or not? What I said was that for artists who are complaining about their copyright being abused , it does not matter. 10 years ago they were not complaining, because AIs looking like ChatGPT (to users who see it as a black box) did not exist (or were not remotely as powerful). And I understand that. It is not their job to learn how the black box wo…

Because fair use allows transformation and the output of their algorithm looks nothing like the input of the copyrighted work? For generative models its more complicated because generative models can actually reproduce large sections of a copyrighted work so transformation is a bit less clear.
Post reply on HN