Live data from Hacker News

Authors say OpenAI 'ingested' their books to train ChatGPT

businessinsider.com

11–20 of 49 posts

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#11

Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?

An LLM isn’t a human brain and it’s almost always a mistake to think of it as one.

The LLM is more akin to a search engine with a much more advanced interface. If Google published the entirety of a book without having the rights they’d be taken to court too.

Don’t get me wrong I’m not saying they’re the same, I think there are convincing arguments to be made either way. I just don’t think the comparison to a human brain is the right one.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#12

Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?

Pretty sure you answered that in your response.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#13
post #9
post #8

> They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology. If ChatGPT can reliably answer "what happens on page (x) of book (y)?" that would provide fairly convincing evidence, as summaries or study notes on a random book are unlikely to be that detailed. Moreover, if this approach worked consistently, it would enable…

not sure i agree with your conclusion since LLMs can easily avoid mentioning the page even if they knew it. Though id also guess that the llm didnt ingest page numbers and just ingested books raw? Anyway i tried it. > im looking king for a quote in the book iRobot, what page is isaac asimovs the laws of robotics on ? provide the version of the book so as well so were clear ChatGPT I'm sorry, but as an AI text-based m…

> it doesnt know

Correct, because it only knows the text, not the layout. It will tell you that it appears in Runaround, and that the Laws are "usually presented towards the beginning of the story, in the dialogue between the characters Powell and Donovan".

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#14

Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?

For one thing, a single humans knowledge cannot be scaled to serve hundreds of millions of users simultaneously.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#15

Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?

I still don't see why the person or company who owns the LLM would owe the book author anything more than the price of the book. Its not like the LLM is going to output or reproduce the book in its entirety. You'll ask a relevant question, it will answer based on the data it has, not just whats in that book.

Same as if you ask a human a relevant question about a book they read. Like an LLM, they can't print you a full copy of that book. They can only answer the question based on their knowledge of the book and other things.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#17
It's really clear to me that a new legal framework will be required to deal with the societal consequences of advanced AI. For example, if this article were instead about, say, an actual human who read all of these books, and then posted reviews and some summaries online, and then this human got sued by the authors, I think every single comment here would be decrying this as an abuse of copyright law, and that this should fall squarely under fair use.

The thing that is "scarier" if you will about what AI can do is the sheer speed and breadth that it is capable of. It really is the scale that changes how people feel about these technologies, and that requires new legal frameworks in my opinion because I think people really feel what is "fair use" is different if it comes from a person vs. a machine.

Similar analogy: before the Internet, pretty much everyone agreed you didn't have an expectation of privacy if you were walking around outdoors. But there is a marked difference in thinking "Yeah, I expect other people walking around may see me, or even take a picture of me" compared to "I think someone should be able to take a picture of me and put it on the Internet with my accurate geolocation so the entire planet can look it up, for all time."

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#18

It's really clear to me that a new legal framework will be required to deal with the societal consequences of advanced AI. For example, if this article were instead about, say, an actual human who read all of these books, and then posted reviews and some summaries online, and then this human got sued by the authors, I think every single comment here would be decrying this as an abuse of copyright law, and that this s…

Essentially they want a copyright on knowledge, Such a law can be misused. Also its hard to prove this just by looking at the weights, and for something as large as GPT-3.5 it would not be easy to prove this.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#19
post #9
post #8

> They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology. If ChatGPT can reliably answer "what happens on page (x) of book (y)?" that would provide fairly convincing evidence, as summaries or study notes on a random book are unlikely to be that detailed. Moreover, if this approach worked consistently, it would enable…

not sure i agree with your conclusion since LLMs can easily avoid mentioning the page even if they knew it. Though id also guess that the llm didnt ingest page numbers and just ingested books raw? Anyway i tried it. > im looking king for a quote in the book iRobot, what page is isaac asimovs the laws of robotics on ? provide the version of the book so as well so were clear ChatGPT I'm sorry, but as an AI text-based m…

Ah interesting. That could be avoidance though. Guess it's useless for pseudo-copy protection schemes for some old games like "what's the first word in the third paragraph on page 17 of the game manual."

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#20

It's really clear to me that a new legal framework will be required to deal with the societal consequences of advanced AI. For example, if this article were instead about, say, an actual human who read all of these books, and then posted reviews and some summaries online, and then this human got sued by the authors, I think every single comment here would be decrying this as an abuse of copyright law, and that this s…

Essentially they want a copyright on knowledge, Such a law can be misused. Also its hard to prove this just by looking at the weights, and for something as large as GPT-3.5 it would not be easy to prove this.

You’re right. Knowledge provided by authors isn’t unique and as long as AI doesn’t literally cite complete chapters word for word, this doesn’t seem to have anything to do with copyright.
Post reply on HN