Live data from Hacker News

Fear of AI just killed a useful tool

techdirt.com

291–300 of 309 posts

Re: Fear of AI just killed a useful tool

#291
post #166

Earlier quoted context omitted.

The best response, for us all collectively, is to always ignore everyone's opinion online. There is zero value in anything on reddit, twitter, facebook, the media these days. Just ignore it. All of it. Outrage or not. I see downvotes, but I mean it. You know who you listen to? Your friends. Your neighbours. Your local community. You listen to PEOPLE, not sockpuppets. You listen to legitimate human beings, not AI gene…

> Your neighbours. Your local community. You listen to PEOPLE, not sockpuppets. You listen to legitimate human beings, not AI generated blather, or curated news stories, or groups working together to generate hate, outrage, to stoke anger, upset. You either have a significantly better social circle than I do or are glossing over a bunch of nuance. Some of my family back east have been getting their brains rotted by f…

I said nothing of abstainment. In fact, I am posting!

Ignoring online comments, especially criticism, does not mean disregardment. And note, context is important. Note what I am replying to.

Simply put, on a medium where one person can sockpuppet appear as 1000, where one person can rally 1000 useful idiots with one disingenuous post, one cannot care what is said.

Ignore it.

We already have 30 year old adults, trying to discuss political nuance online, not realising that they may be piled on by a dozen 8 year olds. People presume the person behind the text is real, the person is their approximate age, or at least an adult, that the person is debating in good faith.

None of this is necessarily true, and in any large group of responses, the above chicanery is happening.

No one should care what a bunch of "people" on Twitter say.

Re: Fear of AI just killed a useful tool

#292

Earlier quoted context omitted.

>> had no conceivable way of threatening the original authors, financially or otherwise how so? what's inconceivable about it? >> seemed to hinge entirely on the irrational fear how are the authors "fearing without reason" or "illogically fearing" ?

> how so? what's inconceivable about it? Authors make money through sales of their work. This tool was a writing aid that analyzed text and included some copyrighted works in its dataset. There was no way of retrieving these books in full, and the excerpts that were allegedly shown to the users were used in an analytical context, unlike the original works. So, this website couldn't replace ownership of the actual boo…

>> Basically, the way the service used this data

it's not about the past (2017...), the authors are concerned about how the dataset could be or is likely to be used from now on.

many tech projects these days have or are thinking about integrating third-party A.I. providers in their services, either to harness the power of their large datasets or their large user-base. I think it's great if authors/users opt-in to this, but likewise I agree with those that want out (opt-out).

>> I called it "fear" because there was no strong argument on the authors' side as to why this tool is bad

their argument doesn't have to be a peer-reviewed journal, it suffices to say "i don't want my books in your dataset"

>> I called it "illogical" because I think that it's no coincidence that this controversy only came up now, in 2023. Back in 2017 and onward, the existence of this tool didn't appear to generate pushback ...

6 years have passed since 2017, life moves on, it's natural for things to change e.g: the project's code, updating servers, partnerships, emergence of third-party tools/libs/services etc etc.

Re: Fear of AI just killed a useful tool

#293
post #192

Earlier quoted context omitted.

Sure. I really did not mean that specifically for Prosecraft. But the article questions why authors are attacking Prosecraft "because it does no harm". My answer is that authors don't (and can't, really) make the difference in a per-case basis. At this point what they see is that LLMs trained on their copyrighted material are able to generate similar material thanks to their copyrighted material that was used in the…

I think most people who have thought about it understand the impact AI models seem destined to have on writing (and digital 2d art, soon music, and later other things). In addition to writers and voice artists panicking, see the Hollywood strikes, for instance, and what's currently happening in the corporate world to digital artists. Copyright is not the correct tool to address it. In the U.S., the basis for copyrigh…

> I think most people who have thought about it understand the impact AI models seem destined to have on writing

Go back to the beginning of social media, and tell me that "most people who had thought about it had understood the impact social media would have on society". It is really not a given. And that is my criticism: we see from history that it is not straightforward to understand the impact of new technology, but we engineers keep making the same mistakes over and over again.

> Copyright is not the correct tool to address it

Maybe not, that's right. I don't think anyone disagrees. The issue - at least from the point of view of artists - is more that some people (including authors and artists) want the problem addressed, and others (including engineers) just want to make money with their new toy and don't care much about addressing the problem.

> doesn't that imply that AI is better at generating useful entertainment than humans are?

I don't think so, no. It is maybe economically more successful, but I think it is clear that what is good for the economy is not necessarily good for society.

> however, it applies to fiction content first

Well... that is ignoring all the black hat use-cases, going from phishing to political mass manipulation, I would say :-)

Re: Fear of AI just killed a useful tool

#294
post #191

Earlier quoted context omitted.

> When an actual human author mimics other writers styles, it is not illegal, why exactly should it be illegal for an author to use an LLM to do it? There is a fundamental difference of scale. Say I write a blog post about some technical thing I know. You read it, learn from it (and other sources), and then you write your own blog post with your understanding. You may link to my post (if you believe it is heavily ins…

I do not disagree with any of what you wrote. That is also an entirely different line of reasoning than your first argument. That said, LLMs today, cannot do this in a meaningful way. If an author cannot write a better book than ChatGPT, then that author would not be able to live of their writing anyway. And the authors that use ChatGPT to write a book, but still put the effort into fine tuning it, will not be able t…

> We cannot prevent it from being used for training, so the next logical step is to protect it the same way as we protect technology with patents.

Why not? Just say that using for training is considered derivative work, and that's it. Now copyright owners just have to update their license to allow for training if they want to, and that's solved. Of course, Big Tech makes less money from that scenario.

> I am certain that trying to ban LLMs, or dictating what and how is not the answer.

I wouldn't ban LLMs because of copyright issues, though I would let authors choose whether their IP can be used for training or not.

However, copyright is only one issue with LLMs. All the black hat use-cases are a whole other category of issues. And I am of the opinion that technology is not neutral: IMO, it is perfectly fine for a society to ban a technology if it believes that it is globally doing more harm than good.

Re: Fear of AI just killed a useful tool

#295
post #278

Earlier quoted context omitted.

I think it is not completely off topic. Here is how I see it: Engineers tend to globally think that LLMs are not really a problem for copyright holders. At least those who develop LLMs pretty clearly don't give a damn. And on top of that, it is in their interest to not be constrained by copyrights. If this is my feeling (that engineers globally don't care about copyright holders), then it seems reasonable to me that…

To rephrase in my own understanding of what you wrote: 1) Some engineers (or more broadly, software developers) do not respect copyright 2) Therefore you reasonably are skeptical of projects related to material under copyright. 3) It is not always obvious if a project is respectful of copyright. Now, applying these #1,#2,#3 you believe they justify the outrage for this particular project. I disagree, because outrage…

> you believe they justify the outrage for this particular project.

No, I believe it explains it.

> it will make the dismissiveness you predict a self-fulfilling prophecy.

That's the thing: both parties need to listen to each other. The problem here is not this particular project, but the fact that we are not addressing the bigger concern which is LLMs.

IMHO, it is completely useless to try to solve this particular case, because it will happen over and over again. We need to address the LLM issue.

Re: Fear of AI just killed a useful tool

#296
post #119

I am a bit confused about what's so outrageous about this tool. It seems that both the book authors, and some of the people in the discussion here, conflate rudimentary statistics about a book (number of words of certain kind) with the latest wave of generative AI. They are very different in both what value they provide, and what risk they pose to book authors. The tool that book authors got outraged about only provi…

If you read through the angry Twitter thread it's clear that almost everyone thinks that either a) the site is a pirate site that lets you download books or b) that the site lets you generate works in the style of an author. Neither of which is true of course. There are a handful (like I think it's clear though that most of the outrage would still be there even if the author had purchased each and every book.

> I actually don't know about the legality of something like that.

Techdirt's analysis of the legality seems correct to me. TL;DR is that it seems legal.

Re: Fear of AI just killed a useful tool

#297
post #123
post #67

If you want to do this kind of thing, let authors opt-in (or publishers). Yes, it will take effort and probably go slow, but if the tool is really useful and amazing, it should be doable. I suspect the authors are put-off by a couple things: - the text of the works scanned seems like it may be from pirated sources. That poisons the project, no matter what it does with the scans, for many authors. - the use of these s…

> If you want to do this kind of thing, let authors opt-in (or publishers). If it's fair use, why should you have to do that? The same copyright law protecting author's ownership rights over their art also provide "fair use" to other people. Someone may disagree with current fair use laws (and I suspect many outraged here do not), but that's a broader issue not related to this particular tool. It just 100% seems like…

> If it's fair use, why should you have to do that?

You may not be legally required to do that, but it can be an excellent move that benefits you nonetheless.

Much like how Weird Al isn't legally required to get permission to make a parody of a popular song, but he does so anyway.

But in this case, I don't think you even need to invoke Fair Use. I think what he did simply isn't a copyright violation in the first place.

In reality, the legality of this was never the issue anyway. The issue was that doing this made the authors angry, and the dev didn't want that.

Re: Fear of AI just killed a useful tool

#298

Earlier quoted context omitted.

> how so? what's inconceivable about it? Authors make money through sales of their work. This tool was a writing aid that analyzed text and included some copyrighted works in its dataset. There was no way of retrieving these books in full, and the excerpts that were allegedly shown to the users were used in an analytical context, unlike the original works. So, this website couldn't replace ownership of the actual boo…

>> Basically, the way the service used this data it's not about the past (2017...), the authors are concerned about how the dataset could be or is likely to be used from now on. many tech projects these days have or are thinking about integrating third-party A.I. providers in their services, either to harness the power of their large datasets or their large user-base. I think it's great if authors/users opt-in to thi…

> the authors are concerned about how the dataset could be or is likely to be used from now on.

The service in question wasn't introducing anything new that'd appear to justify all the recent pushback. It feels like you continue generalizing your statements, while the discussion topic is about what made Prosecraft specifically so preposterous that it warranted the outrage.

> their argument doesn't have to be a peer-reviewed journal, it suffices to say "i don't want my books in your dataset"

It's kind of a blunt statement, but why should they have a say? For example, say I create a website where I publish technical analysis of famous literary works, including basic statistics about a book and a review. Should the authors be able to just take that down? This use is legally protected (as is creating a dataset), so allowing authors to restrict this use seems as arbitrary as allowing them to say that no person can ever bring their books into the country of Moldova or that no one over the age of 50 may read it.

> it's natural for things to change e.g: the project's code, updating servers, partnerships, emergence of third-party tools/libs/services etc

And yet, in this specific situation, all of this is conjecture. Nothing about the project changed in some significant way in 2023 that would warrant this. Further proving it is that the people that are against Prosecraft don't seem to bring up any specific changes or reasons for their stance, only that it is "AI".

Re: Fear of AI just killed a useful tool

#299

Earlier quoted context omitted.

>> Basically, the way the service used this data it's not about the past (2017...), the authors are concerned about how the dataset could be or is likely to be used from now on. many tech projects these days have or are thinking about integrating third-party A.I. providers in their services, either to harness the power of their large datasets or their large user-base. I think it's great if authors/users opt-in to thi…

> the authors are concerned about how the dataset could be or is likely to be used from now on. The service in question wasn't introducing anything new that'd appear to justify all the recent pushback. It feels like you continue generalizing your statements, while the discussion topic is about what made Prosecraft specifically so preposterous that it warranted the outrage. > their argument doesn't have to be a peer-r…

as I said earlier "6 years passed" since 2017 and the arrival of A.I. is causing authors (to say the least) real concern. the industry/ecosystem around X is enough to change to affect those associated with X.

you offered a "pet theory" (as you say) further up, but not much by way of proof that their fear is irrational.

i think we've reached in impasse here, the points are re-cycling.

Re: Fear of AI just killed a useful tool

#300
post #277
post #230

Earlier quoted context omitted.

Because fair use allows transformation and the output of their algorithm looks nothing like the input of the copyrighted work? For generative models its more complicated because generative models can actually reproduce large sections of a copyrighted work so transformation is a bit less clear.

> Because fair use allows transformation and the output of their algorithm looks nothing like the input of the copyrighted work? I feel like you miss the point of a law. You seem to read the law, and say "well, the law says X, new technology Y is compatible with it, so that's legal, everyone is happy". But that is wrong. The law reflects the society we want. Do we want a society that completely kills creative work be…

I think the actual issue is nuanced and complicated. I think it's fairly clear the tool in question that was non-generative AI is the kind of thing we want to allow under fair use. Whether we want to allow generative AI is more complex, I'd lean towards requiring a license because of non-deterministic duplication. Fair use is an important part of copyright law and we should be very cautious about eroding it. For example, I like Green Day's transformation of the scream icon and think it was substantially different enough that it should be allowed. The courts agreed under current transformation laws but if we weaken protection against transformation we likely reverse the ruling of cases like that as well.
Post reply on HN