Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

91–100 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#91
post #3

I think the headline is a bit misleading. Mets did pirate the works but may be entitled to use them under fair use. It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue…

The problem is that "harm" as defined by copyright law is strictly limited to loss of sales due to breach of that copyright; it makes no allowment (that I know of) to livelihoods lost by the theft of the work indefinitely, as AI boosters suggest their tools can do (replace people). The way this court case is going, it's an uphill battle for the plaintiffs to prove concrete harm in that very narrow context, when the r…

A difficult, but not intractable problem: OLMoTrace claims to be able to trace from output to training data in seconds [1]. Notably, it can do this because OLMo itself was intentionally designed to be open and transparent [2]; it was trained on 4.6 trillion tokens of entirely open data (which you can download yourself) [3]. There's nothing stopping Meta or OpenAI from creating a similar tool, other than the obvious detail of that showing their exact training data.

[1] https://arxiv.org/abs/2504.07096

[2] https://allenai.org/blog/olmotrace

[3] https://huggingface.co/datasets/allenai/olmo-mix-1124

Re: Judge said Meta illegally used books to build its AI

#92

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

Copyright is defined in law and as the original poster stated, whether this is 'copying' as defined by copyright law is legally ambiguous.

Copyright doesn't protect against all forms of duplication. For instance, you own the copyright to your post and grant HN a license to offer copies of it. I have no direct license from you to copy the content of your post; but I can copy it to memory, copy a cache to disk, and make a copy appear on my display.

Re: Judge said Meta illegally used books to build its AI

#93

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

>That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

The NYTimes in 2023 was able to demonstrate that the models can reproduce entire articles verbatim[0] with minimal coercion.

[0]https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

Re: Judge said Meta illegally used books to build its AI

#94

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

If you read a book and later understand its plot but can only explain it in your own words, did you copy it? The model isn’t storing the book.

The model doesn't "understand its plot". So I am not sure this is a good analogy.

Re: Judge said Meta illegally used books to build its AI

#95
post #54

Earlier quoted context omitted.

Copyright was invented (in its modern form) by corporations. It will be uninvented if need be for corporations.

I'm curious what you mean by "in it's modern form". You seem to suggest there was a previous form that was not invented by corporations, but I don't believe that is the case.

There were two, major differences in prior forms of copyright:

1. It protected works to reward authors during their lifetime. This was changed to lasting a long time after the author was dead. Then, also for corporations that were only persons on paper and theoretically immortal. This shift let companies squeeze money out of monopolized ideas for over a century rather than supporting artists and their creations. Instead of supporting the small fish, copyright law can reinforce the dominance of the sharks and whales.

2. Copyright was shorter in the U.S. at 28 years with possible renewal. That would balance two goals: give author time to make money off the work; let society use the work in a timeframe where it would still matter to them. Now, we can't have most works until long after they're useful in the market. We might not even speak the language they spoke, like older vs current English.

Personally, I'd love to see a limit of 5-20 years on copyrighted works. If authors want more money, they can make more stuff. Allowing remixes of culturally and technologically relevant content will create huge, thriving ecosystems. I think my concept is also proven out by the open source ecosystem.

A limit would also be great for legal AI. We could train them on all human content up to 5-20 years ago. Tons of jobs would be created digitizing and optimizing that content. Then, companies would pay to create or license modern content that updated those foundation models. Under current law, it would be impossible for smaller companies to build highly-competitive A.I.'s due to licensing cost and arbitrary restrictions.

Re: Judge said Meta illegally used books to build its AI

#96
post #55

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

I'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did no…

The original complaint alleges that the training process requires copying the material into the model and thus requires consent of the copyright holder. (Copyright protects copying but notably not use, so the complaint has to say they copied it in order to have standing). Then it says they didn't have consent.

They also mention Books3, but they don't appear to actually allege anything against Meta in regards to it and are just providing context.

I don't think it actually changes anything material about this complaint if Meta bought all the books at a bookstore since that also doesn't give you the right to copy the works.

The original complaint is 2 years old though, so I don't really know the current state of argumentation.

https://www.courtlistener.com/docket/67569326/1/kadrey-v-met...

Note that incidental copying (i.e. temporary copies made by computers in order to perform otherwise legal actions) is generally legal, so "copying" in the complaint can't refer merely to this and must refer more broadly to the model itself being a copy in order to have standing.

Re: Judge said Meta illegally used books to build its AI

#97

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

There’s something called a substantive transformation test in copyright law. When you write a summary of a book, you don’t infringe on copyright because it’s a “substantial transformation”. This goes along with the idea that you can copyright the text but not the ideas it expresses.

When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong argument that it is.

Re: Judge said Meta illegally used books to build its AI

#98

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

If you read a book and later understand its plot but can only explain it in your own words, did you copy it? The model isn’t storing the book.

> If you read a book and later understand its plot but can only explain it in your own words, did you copy it?

I think that is the center of the conversation. What does it mean for a computer to "understand"? If I wrote some code that somehow transformed the text of the book and never return the verbatim text but somehow modified the output, I would likely not be spared because the ruling will likely be my transformation is "trivial".

Personally, I think we have several fixes we need to make:

  1. Abolish the CFAA. 
  2. Limit copyright to a maximum of 5 years from date of production with no extension possible for any reason. 
  3. Allow explicit carveout in copyright for transformational work. Explicitly allow format shifting, time shifting, yada yada. 
  4. Prohibit authors and publishers from including the now obviously false statements like "No part of this publication may be reproduced, distributed, or transmitted in any form or by any means, including photocopying, recording" bla bla bla in their works. 
  5. I am sure I am missing some stuff here.
For brand protection, we already have trademark law. Most readers here already know this but We really should severe the artificial ties we have created between patents, trademarks, and copyright.

https://www.gnu.org/philosophy/not-ipr.en.html

Re: Judge said Meta illegally used books to build its AI

#99

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

If you read a book and later understand its plot but can only explain it in your own words, did you copy it? The model isn’t storing the book.

I just happen do read the Phoenix Technologies wikipedia page a few days ago. This company is known for developing BIOS software for computers. Maybe you've seen their logo when you first turn on your computer.

In early computing, everything was closed sourced. Quoting the wikiepdia page,

To develop a legal BIOS, Phoenix used a clean room design. Engineers read the BIOS source listings in the IBM PC Technical Reference Manual. They wrote technical specifications for the BIOS APIs for a single, separate engineer—one with experience programming the Texas Instruments TMS9900, not the Intel 8088 or 8086—who had not been exposed to IBM BIOS source code.

The legal team at Phoenix deemed inappropriate to "recall source in their own words" for legal reasons.

My non-legal intuition is that these companies training their models are violating copyright. But, the stakes are too high--it's too big to fail if you will. If we don't do it, then our competitors will destroy us. How do you reconcile that?

Re: Judge said Meta illegally used books to build its AI

#100

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

> AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling.

IANAL, but it doesn't look that hard. On first glance this is a fair use issue.

What an LLM spits out is pretty clearly transformative use. But the fact that it pulls not only the entirety of the work, but the entirety of MOST works means that the amount is way beyond what could be fair use. Plus it's commercial use. Put it together and all LLMs are way illegal.

Post reply on HN