Live data from Hacker News

Authors say OpenAI 'ingested' their books to train ChatGPT

businessinsider.com

21–30 of 49 posts

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#21
post #11

Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?

An LLM isn’t a human brain and it’s almost always a mistake to think of it as one. The LLM is more akin to a search engine with a much more advanced interface. If Google published the entirety of a book without having the rights they’d be taken to court too. Don’t get me wrong I’m not saying they’re the same, I think there are convincing arguments to be made either way. I just don’t think the comparison to a human br…

which, for anybody following along at home, is related to something that happened back in 2015.

Google was trying to digitize the world's books and got sued for it.

https://www.npr.org/sections/thetwo-way/2015/10/16/449172748...

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#22
>They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology.

This just isn't true. The authors have detailed summaries of their books on Wikipedia, that claim seems unsustainable.

There are actually interesting legal questions about AI, but this case seems not that interesting. I don't even see how the authors would demonstrate that their copyrighted works were used and not some summary, sourced from elsewhere.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#24
post #23

Ok. I may as well come clean. For the past 36 years, I have been ingesting all of the information available to train a model called b33-j0r I didn’t mean to read your books, they were mostly trash. I’m so sorry for reading things you published!

I hope this point of view wins eventually. It's the only sensible one.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#25
Fundamentally this is about wealth distribution. The society do want those data to be used for AI. It's just the data providers should be compensated, instead of OpenAI just gatekeeping their product that critically depend on those data. The original model of copyright just doesn't work anymore in this new era.

Music streaming has changed the music industry. Can we expect something similar happening in AI?

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#26
Same old Murican business model. Bill Gates did it, Zuck did it, Jeff did it, Elon does it. Embrace Extend Extinguish, Facebook forced tracking into the web and made profiles for people who didn't sign up, Google absorbed everything, ect.

The American business model has been Robber Barons. Since the 1800s, at least. OpenAI is the latest. The reality is, you can never give anything valuable to the 'free market' and an internet connected device, because you're never going to be able to govern it according to your values.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#27
post #11

Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?

An LLM isn’t a human brain and it’s almost always a mistake to think of it as one. The LLM is more akin to a search engine with a much more advanced interface. If Google published the entirety of a book without having the rights they’d be taken to court too. Don’t get me wrong I’m not saying they’re the same, I think there are convincing arguments to be made either way. I just don’t think the comparison to a human br…

a LLM is closer to a very primitive brain than a search engine. A search engine can't make mistakes, it output data verbatim, it may not be able to find something, but it's never making mistakes. A LLM have an input, and an output, which can wildly differ, it can respond to your demands, even trigger other software with relevant data.

It's way closer to a primitive brain than to a search engine.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#29

>They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology. This just isn't true. The authors have detailed summaries of their books on Wikipedia, that claim seems unsustainable. There are actually interesting legal questions about AI, but this case seems not that interesting. I don't even see how the authors would demon…

It seems relatively straight forward (famous last words) to assess whether actual copyrighted text is embedded within the network. If you can prompt output that includes verbatim extracts when the copyright avoidance post-processing is disabled then you know that it has been consumed.

Of course whether that was purposeful or inadvertently as a part of the larger training set would not be determined but you would know that the text is in there.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#30
There's a lot of discussion about whether existing copyright laws apply to training data, but not enough discussion of what should change in copyright law to give the best outcome to society.

I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new.

It seems like society would benefit from have ML systems around that can be trained on copyrighted material, but not if it prevents authors from being able to make a career out of producing great work. We need to balance the outcome for owners of ML companies, authors, and society as a whole, ideally prioritising the latter.

What do you think we should do?

Post reply on HN