Live data from Hacker News

Pearson taking legal action over use of its textbooks to train language models

standard.co.uk

1–10 of 64 posts

Re: Pearson taking legal action over use of its textbooks to train language models

#2
My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts.

So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material well.

It made me wonder this exact point. If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair?

Edit: I guess this isn't really that different than the fact I read the book and have described it to many friends and even done the exercise with them.

Re: Pearson taking legal action over use of its textbooks to train language models

#3

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair?

What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it less fair for the AI to do what is effectively the same thing? Is it simply a matter of scale? Or is there something more fundamental at work?

Re: Pearson taking legal action over use of its textbooks to train language models

#4

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

That was basically my edit, we may have missed each other.

The only point I can see making is that for some reason I doubt openai actually paid the author for the contents of the book, but maybe they did.

Maybe they borrowed ebooks one by one from some digital library and ingested them that way for free.

Re: Pearson taking legal action over use of its textbooks to train language models

#5

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

> It made me wonder this exact point. If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair?

Dr Burns seems to have done well, so in his case I'd think he was at least adequately compensated for his life's work. You can't answer "fairness" on copyright in all cases just by that, but this particular case I feel there's little or no harm done.

If people start selling these GPT things because they let you avoid fair recompense, that's another thing we'll have to work out as a society.

Re: Pearson taking legal action over use of its textbooks to train language models

#6
> Bird also said it was usually easy to tell what a large language model such as ChatGPT has been trained on, because “you can ask it”.

So what does that mean? They literally asked “Did you read book X during training”?

Sounds like they’re fools. Desperate fools, most likely.

Re: Pearson taking legal action over use of its textbooks to train language models

#7

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

I'd argue ChatGPT is closer to restating copyrighted materials in a different way. And that is legally questionable. Whether that actually matters depends on how much money is at stake and how much each party cares to spend to enforce or protect their viewpoint.

Allowing these systems to utilize copyrighted materials without compensating the copyright holder is a legal loophole the size of the sun. Then it becomes a perfectly legal defense to write "an AI" that slurps up copyrighted material and produces results that are "based on" that material.

My issue is that the copyrighted work is fundamentally integrated into the algorithm. Say you wrote a program that would input a copyrighted work, then prints out the contents, but with some words replaced with synonyms. Would the output of that be okay? What about if phrases were replaced? How about if entire sections are added, rearranged, and/or removed? At some point, the results become so mangled as to not be recognizable, but the fact that a copyrighted work serves as an input to the program is still an issue, IMHO.

A major difference between humans and computers is that a human can't (generally) recall hundreds of pages of text and regurgitate it on demand. So a human who has read a book only retains a fraction of the contents, and even then, generally they don't retain all of the information accurately.

Re: Pearson taking legal action over use of its textbooks to train language models

#8

Earlier quoted context omitted.

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

That was basically my edit, we may have missed each other. The only point I can see making is that for some reason I doubt openai actually paid the author for the contents of the book, but maybe they did. Maybe they borrowed ebooks one by one from some digital library and ingested them that way for free.

Somehow I feel like they just scraped libgen and sci hub

Much more "scalable" i.e. easy to get raw training dataset. Move fast and break things

Re: Pearson taking legal action over use of its textbooks to train language models

#9

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

I do wonder if it will disincentivise the author from writing the book in the first place. What’s the point if an AI is going to regurgitate the material to every Tom, Dick and Harry thereby killing the demand for the book.

Re: Pearson taking legal action over use of its textbooks to train language models

#10

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

Answer is pretty obvious: One is a computer program, literally designed to pilfer copyrighted works, the other is a human being.
Post reply on HN