Live data from Hacker News

Authors say OpenAI 'ingested' their books to train ChatGPT

businessinsider.com

31–40 of 49 posts

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#31

There's a lot of discussion about whether existing copyright laws apply to training data, but not enough discussion of what should change in copyright law to give the best outcome to society. I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new. It seems like society would benefit from have ML s…

Considering that I'm not allowed to copy any of the physical books, films or songs that I own as it breaks copyright then copying it and putting it into a generative AI should also be illegal. Just because you change it from words into weights, it's just another form of compression. I'd rather not live in the world where AI models have better rights than people.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#32
post #23

Ok. I may as well come clean. For the past 36 years, I have been ingesting all of the information available to train a model called b33-j0r I didn’t mean to read your books, they were mostly trash. I’m so sorry for reading things you published!

Presumably you paid for some of the trash you read, either through buying a book outright, borrowing from your library (taxes), etc?

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#33
post #9

Earlier quoted context omitted.

not sure i agree with your conclusion since LLMs can easily avoid mentioning the page even if they knew it. Though id also guess that the llm didnt ingest page numbers and just ingested books raw? Anyway i tried it. > im looking king for a quote in the book iRobot, what page is isaac asimovs the laws of robotics on ? provide the version of the book so as well so were clear ChatGPT I'm sorry, but as an AI text-based m…

Ah interesting. That could be avoidance though. Guess it's useless for pseudo-copy protection schemes for some old games like "what's the first word in the third paragraph on page 17 of the game manual."

fwiw ive asked it to summarize the first 5 chapters before, and it was able to do that. Just a high level breakdown

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#34
post #29

>They argue that ChatGPT's ability to produce detailed summaries of their works indicates their books were included in datasets used to train the technology. This just isn't true. The authors have detailed summaries of their books on Wikipedia, that claim seems unsustainable. There are actually interesting legal questions about AI, but this case seems not that interesting. I don't even see how the authors would demon…

It seems relatively straight forward (famous last words) to assess whether actual copyrighted text is embedded within the network. If you can prompt output that includes verbatim extracts when the copyright avoidance post-processing is disabled then you know that it has been consumed. Of course whether that was purposeful or inadvertently as a part of the larger training set would not be determined but you would know…

>If you can prompt output that includes verbatim extracts

If I create a program that picks random words from a dictionary and I end up with a seed that generates that text verbatim, then does that mean my program contains the copyrighted text?

You might be able to craft an intricate prompt that just happens to recreate that copyrighted text. Run it enough times until you get it verbatim and done.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#35

There's a lot of discussion about whether existing copyright laws apply to training data, but not enough discussion of what should change in copyright law to give the best outcome to society. I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new. It seems like society would benefit from have ML s…

Considering that I'm not allowed to copy any of the physical books, films or songs that I own as it breaks copyright then copying it and putting it into a generative AI should also be illegal. Just because you change it from words into weights, it's just another form of compression. I'd rather not live in the world where AI models have better rights than people.

>Considering that I'm not allowed to copy any of the physical books, films or songs that I own

You should be able to and you can in some/many countries.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#36

There's a lot of discussion about whether existing copyright laws apply to training data, but not enough discussion of what should change in copyright law to give the best outcome to society. I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new. It seems like society would benefit from have ML s…

The same thing we did when basket weavers were upset robotic manufacturing took over. Let supply and demand decide what society considers valuable. If we want the $2 basket that gets delivered yesterday but breaks in 1 year, so be it. If the master basketweaver thinks there is a market for selling human made baskets at $100 a pop, but they last 10 years, then prove it.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#37

There's a lot of discussion about whether existing copyright laws apply to training data, but not enough discussion of what should change in copyright law to give the best outcome to society. I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new. It seems like society would benefit from have ML s…

>but not if it prevents authors from being able to make a career out of producing great work

Why? Why should we hold back technological and social progress because it might make it more difficult for some authors to have a profitable career?

Are only people that have the mechanical skill to draw and write well the only ones that should be allowed to profit from creative work? Because that's how it's been so far. You might have the best ideas for a story, picture or movie, but if you lack the mechanical skill to bring them to life then your creativity is worthless. AI generation can help those people have a chance too.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#38

What are we going to class these AI's as though? If I read a book and then tell someone else about it, that isn't copyright infringement, if the AI 'reads' the book and tells someone else about it, they are claiming it is?

Here's an idea for a dystopian sci fi novel: AI learning is treated differently from human learning, therefore we start building computers with human/animal brains.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#39

Same old Murican business model. Bill Gates did it, Zuck did it, Jeff did it, Elon does it. Embrace Extend Extinguish, Facebook forced tracking into the web and made profiles for people who didn't sign up, Google absorbed everything, ect. The American business model has been Robber Barons. Since the 1800s, at least. OpenAI is the latest. The reality is, you can never give anything valuable to the 'free market' and an…

The problem is that this isn't only limited to OpenAI though. The same rules will also bind efforts like StabilityAI that does give away their models for free for everyone to use.

Re: Authors say OpenAI 'ingested' their books to train ChatGPT

#40

There's a lot of discussion about whether existing copyright laws apply to training data, but not enough discussion of what should change in copyright law to give the best outcome to society. I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new. It seems like society would benefit from have ML s…

Considering that I'm not allowed to copy any of the physical books, films or songs that I own as it breaks copyright then copying it and putting it into a generative AI should also be illegal. Just because you change it from words into weights, it's just another form of compression. I'd rather not live in the world where AI models have better rights than people.

Can you point to any substantial reproductions of copyrighted novels that are output by LLMs? If it’s just compression, you should be able to pull a substantial reproduction of the works out of them.
Post reply on HN