Live data from Hacker News

Pearson taking legal action over use of its textbooks to train language models

standard.co.uk

11–20 of 64 posts

Re: Pearson taking legal action over use of its textbooks to train language models

#11

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

> answer questions for people without having to buy the book.. is that really fair?

Yes.

It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothing to the author.

It's also a fundamental principle of copyright law, at least in the United States.

https://supreme.justia.com/cases/federal/us/499/340/

https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R....

Re: Pearson taking legal action over use of its textbooks to train language models

#12

Earlier quoted context omitted.

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

Answer is pretty obvious: One is a computer program, literally designed to pilfer copyrighted works, the other is a human being.

If you smuggle "pilfer" into the description of AI, then in that case it's "obvious".

But the comparison could just as easily be written, "One is a computer program, literally designed to learn from copyrighted works, the other is a human being that learns from copyrighted works"

If ChatGPT is parroting entire paragraphs of copyrighted work in its answers, then "pilfer" is probably the right word. Or maybe "plagiarize". But if it's training on the information gleaned from the work, and using that training to synthesize the answers, isn't that a lot more like how a learned person applies their knowledge?

I haven't yet made up my mind completely on how IP should interact with AI. I see good arguments for and against. But the answer is far from obvious!

Re: Pearson taking legal action over use of its textbooks to train language models

#13

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

AI and humans aren't the same thing. Even if we do anthropomorphize AI.

Re: Pearson taking legal action over use of its textbooks to train language models

#14

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

> No author has the right to prevent me from discussing a book with someone.

a lossy compression function is not "someone"

Re: Pearson taking legal action over use of its textbooks to train language models

#15
post #14

Earlier quoted context omitted.

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

> No author has the right to prevent me from discussing a book with someone. a lossy compression function is not "someone"

[deleted]

Re: Pearson taking legal action over use of its textbooks to train language models

#16

Earlier quoted context omitted.

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

Answer is pretty obvious: One is a computer program, literally designed to pilfer copyrighted works, the other is a human being.

Why is it "pilfering" when the computer program does it, but "learning" when a human does it?

Re: Pearson taking legal action over use of its textbooks to train language models

#17

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

At what point is it discussing a subject, vs basically verbatim listing things, methods, reasoning from a book?

Re: Pearson taking legal action over use of its textbooks to train language models

#18

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

As a side note, how would we all feel if someone wrote a depression /cognitive behavioral therapy ai chat bot, charged $20 a month for it, and then it just basically ran you through the exact methods from the feeling good book?

Re: Pearson taking legal action over use of its textbooks to train language models

#19

Earlier quoted context omitted.

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

I'd argue ChatGPT is closer to restating copyrighted materials in a different way. And that is legally questionable. Whether that actually matters depends on how much money is at stake and how much each party cares to spend to enforce or protect their viewpoint. Allowing these systems to utilize copyrighted materials without compensating the copyright holder is a legal loophole the size of the sun. Then it becomes a…

I get what you're saying, but I'm still not convinced that introducing the notion of "a computer program" vs "a human" changes anything in the equation. That is to say, IF you - as a human - had a "photographic memory" and after reading a copyrighted work began to spit out big chunks of it verbatim in response to a query, then you would already be guilty of copyright infringement, no? Likewise, if a computer program reads 10 books on, I dunno, let's say Quantum Physics, and then - in response to a query - emits a technically correct answer that isn't obviously cribbed from any one book... has it violated copyright?

My point is not to say that contemporary "AI" systems never emit things that are arguably copyright infringement. It's more just to say that there's nothing about being an AI that makes it necessary that its output be copyright infringing.

Re: Pearson taking legal action over use of its textbooks to train language models

#20

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

As a side note, how would we all feel if someone wrote a depression /cognitive behavioral therapy ai chat bot, charged $20 a month for it, and then it just basically ran you through the exact methods from the feeling good book?

This happens all the time with “self help” tiktokers and YouTubers and whatnot that are just recycling others ideas and content.

How would you like to stop this? It doesn’t violate copyright unless they read the book verbatim. I guess ideas could be patented as a utility method by the author and it would prevent someone else from doing that method for 20 years.

Mostly this is just the way things were designed and I think that it’s good in the sense that the purpose of creation is to further human experience, not maximize creator revenue.

Post reply on HN