Live data from Hacker News

Authors say OpenAI 'ingested' their books to train ChatGPT

businessinsider.com

1–10 of 49 posts

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#3

[flagged]

If you plagiarize it, yes, they absolutely are.

They're suing because OpenAI produced summaries of the works without any copyright information.

I'm not sure at what point a summary becomes a plagiarized abridgement, but that's why we have a court system.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#6

[flagged]

If you write a book where Harry, Ron and Hermione go to a magic school you will get sued. If you change the names, perhaps you can claim it's unrelated.

Only if you try to sell it... Otherwise its legally protected.

https://archiveofourown.org/admin_posts/5857

"Free use" legality of these models is what interests me the most. For instance, what if someone releases an excellent model specifically for producing entire Harry Potter stories?

What if the entire training dataset was reworded by an LLM so no one can prove it contains any plagiarized material?

Is the trainer going to get sued? What about the users? Would the entire model be "illegal" like a pirated book even though it ostensibly doesn't contain the book?

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#8
> They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology.

If ChatGPT can reliably answer "what happens on page (x) of book (y)?" that would provide fairly convincing evidence, as summaries or study notes on a random book are unlikely to be that detailed. Moreover, if this approach worked consistently, it would enable the entire book to be summarized - or even rewritten - page by page.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#9
post #8

> They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology. If ChatGPT can reliably answer "what happens on page (x) of book (y)?" that would provide fairly convincing evidence, as summaries or study notes on a random book are unlikely to be that detailed. Moreover, if this approach worked consistently, it would enable…

not sure i agree with your conclusion since LLMs can easily avoid mentioning the page even if they knew it. Though id also guess that the llm didnt ingest page numbers and just ingested books raw? Anyway i tried it.

> im looking king for a quote in the book iRobot, what page is isaac asimovs the laws of robotics on ? provide the version of the book so as well so were clear

ChatGPT I'm sorry, but as an AI text-based model, I don't have direct access to specific book editions or their page numbers. However, I can provide you with general information about Isaac Asimov's "I, Robot" and the Laws of Robotics.

it doesnt know, but i also didnt coax it or jailbreak it to circumvent any avoidance

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#10
Well, to become an expert in any topic, humans have to go to university and read textbooks as a part of the course. Do these humans owe anything to the authors, other than the purchase price of the books? Why should it be any different for AI LLMs/reasoning engines?
Post reply on HN