Live data from Hacker News

Pearson taking legal action over use of its textbooks to train language models

standard.co.uk

21–30 of 64 posts

Re: Pearson taking legal action over use of its textbooks to train language models

#21

Earlier quoted context omitted.

Answer is pretty obvious: One is a computer program, literally designed to pilfer copyrighted works, the other is a human being.

Why is it "pilfering" when the computer program does it, but "learning" when a human does it?

Because those two things are not even equivalent.

A human can reason, make decisions, act in best interest, judge context of a situation, apply it to new scenarios, judge emotion, and create a new derivative work.

The computer program is literally taking someones work and adding fancy way to search through the other person's work, without adding any value.

We need to stop fantasizing these fancy toys are actually any sort of intelligence.

Re: Pearson taking legal action over use of its textbooks to train language models

#22

Earlier quoted context omitted.

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

At what point is it discussing a subject, vs basically verbatim listing things, methods, reasoning from a book?

Reproduction of the precise form of expression is the only thing covered by copyright.

Facts are not protected by any IP law in the United States.

Ideas may be protected by patent law, if the idea constitutes a novel, non-obvious invention.

Re: Pearson taking legal action over use of its textbooks to train language models

#23

Earlier quoted context omitted.

Why is it "pilfering" when the computer program does it, but "learning" when a human does it?

Because those two things are not even equivalent. A human can reason, make decisions, act in best interest, judge context of a situation, apply it to new scenarios, judge emotion, and create a new derivative work. The computer program is literally taking someones work and adding fancy way to search through the other person's work, without adding any value. We need to stop fantasizing these fancy toys are actually any…

We need to stop fantasizing these fancy toys are actually any sort of intelligence.

That strikes me as a rather extraordinary claim. It seems obvious to me that contemporary AI's are "some" sort of intelligence, albeit with open questions around "what kind of intelligence?" and "how intelligent are they?".

Re: Pearson taking legal action over use of its textbooks to train language models

#24

Earlier quoted context omitted.

Why is it "pilfering" when the computer program does it, but "learning" when a human does it?

Because those two things are not even equivalent. A human can reason, make decisions, act in best interest, judge context of a situation, apply it to new scenarios, judge emotion, and create a new derivative work. The computer program is literally taking someones work and adding fancy way to search through the other person's work, without adding any value. We need to stop fantasizing these fancy toys are actually any…

[dead]

Re: Pearson taking legal action over use of its textbooks to train language models

#25

Earlier quoted context omitted.

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

Answer is pretty obvious: One is a computer program, literally designed to pilfer copyrighted works, the other is a human being.

So we can paint ourselves into corner with legislation and lawyers and prevent "AI" progress or have a constitutional referendum to decide if it's fair or not fair use.

Re: Pearson taking legal action over use of its textbooks to train language models

#26

Earlier quoted context omitted.

If you wrote a book on a very rare subject that you were basically the only popular source of detailed info on... And now these models just answer questions for people without having to buy the book.. is that really fair? What if a person bought a copy of the book, read it and learned the material, and then provided advice to people based on the material. I think most people would say that that is fair. So why is it…

I'd argue ChatGPT is closer to restating copyrighted materials in a different way. And that is legally questionable. Whether that actually matters depends on how much money is at stake and how much each party cares to spend to enforce or protect their viewpoint. Allowing these systems to utilize copyrighted materials without compensating the copyright holder is a legal loophole the size of the sun. Then it becomes a…

As far as I can tell, restating materials in a different way does not infringe copyright.

Re: Pearson taking legal action over use of its textbooks to train language models

#27

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

As a side note, how would we all feel if someone wrote a depression /cognitive behavioral therapy ai chat bot, charged $20 a month for it, and then it just basically ran you through the exact methods from the feeling good book?

That was the business model of Joyable, except it charged more and had a human periodically sanity-check your progress on the exercises.

Re: Pearson taking legal action over use of its textbooks to train language models

#28

Earlier quoted context omitted.

I'd argue ChatGPT is closer to restating copyrighted materials in a different way. And that is legally questionable. Whether that actually matters depends on how much money is at stake and how much each party cares to spend to enforce or protect their viewpoint. Allowing these systems to utilize copyrighted materials without compensating the copyright holder is a legal loophole the size of the sun. Then it becomes a…

As far as I can tell, restating materials in a different way does not infringe copyright.

Right, that’s just citing sources. Even if your source is only one book (the Cliff Notes model).

Re: Pearson taking legal action over use of its textbooks to train language models

#29

My favorite anti depression book is Feeling Good by Dr. Burns. I especially like the "if that were true, it would mean that..." exercise in the book, where you drill down to the root cause of your negative thoughts. So I asked chatgpt about the book, and the exercise, and how to do it, and examples of the 9 cognitive distortions you're supposed to label your responses with. It was very helpful and knew the material w…

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

Are you OK with the developers of LLMs capturing the most value of every incremental piece of content created by humans in perpetuity?

The difference is scale. A human can use their learnings from copyrighted work to get a job making $80,000/year. An LLM can use their learnings from copyrighted work to become the biggest, most profitable company ever.

Re: Pearson taking legal action over use of its textbooks to train language models

#30

Earlier quoted context omitted.

> answer questions for people without having to buy the book.. is that really fair? Yes. It's fair because that happens all the time with non-AI intelligences (i.e., people). No author has the right to prevent me from discussing a book with someone. If I were a clinician and incorporated that exercise into my practice, as long as I wasn't copying any content (e.g., worksheets, scoring tables, etc.), then I owe nothin…

Are you OK with the developers of LLMs capturing the most value of every incremental piece of content created by humans in perpetuity? The difference is scale. A human can use their learnings from copyrighted work to get a job making $80,000/year. An LLM can use their learnings from copyrighted work to become the biggest, most profitable company ever.

> Are you OK with the developers of LLMs capturing the most value of every incremental piece of content created by humans in perpetuity?

If it helps humanity as a whole, sure. Artificial gate-keeping based on economic factors such as people being able to keep their $80k jobs is not justifiable to me to stop technological progress; one could have said the exact same thing centuries ago with the advent of any number of inventions.

Post reply on HN