Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

101–110 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#101
post #77

Earlier quoted context omitted.

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

China found the perfect way to disrupt US tech, releasing open source versions of it for free or at least cheaper. Most of US tech is built on open source anyways and with the pace YC is investing in open source alternatives, it will win out in most niches.

My fear is that the US tech won’t be able to compete with state sponsored open source out of China and will move to ban open source or suppress it somehow.

Re: Judge said Meta illegally used books to build its AI

#102
post #94

Earlier quoted context omitted.

If you read a book and later understand its plot but can only explain it in your own words, did you copy it? The model isn’t storing the book.

The model doesn't "understand its plot". So I am not sure this is a good analogy.

To what extent connections in a neural network are analogous to connections between neurons in your brain is open to interpretation and study, but the point of the analogy is that in neither case is a copy being made.

Re: Judge said Meta illegally used books to build its AI

#103

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

That copying is already a violation. At least it was when regular people weee on the receiving end of the lawsuits.

Re: Judge said Meta illegally used books to build its AI

#104

Earlier quoted context omitted.

The output of the LLM is very different from the original, though. It’s hard to look at this and claim it isn’t.

> The output of the LLM is very different from the original, though I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. But I think that this is legally undecided and won't be decided by you…

Yeah, the way the courts decide is unlikely to turn on a detail of how the technology works, so it's difficult for non-legal experts to predict the outcome on the technical merits (since the law has very different priorities).

Music has ended up in a place where short audio snippets are protected by copyright and must be licensed; but for short snippets of text the precedent has generally been that the copying needs to be more substantial. Distributed microplagarism of short phases might end up being ruled to be legal, even if wholesale reproduction is not. Which may not give copyright protection to the generated works, of course, as the question of machine authoring is entirely distinct.

Re: Judge said Meta illegally used books to build its AI

#105

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

The whole point of the fair use clauses is to protect humans. Clearly we can easily say that programs are altogether exempt in favor of humans, and it would be a proper thing to do, until the first real AI is built.

Re: Judge said Meta illegally used books to build its AI

#106

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

The US Federal government operates with the rule that if human eyes don't look at it it doesn't count as a copy or looking at it. This allows them to unconstitutionally spy and log all people's telecommunications. Applying it here it seems pretty clear that corps are within the established bounds. As are any human persons that want to train an LLM this way.

Re: Judge said Meta illegally used books to build its AI

#107
This is the source headline, but it is pure clickbait; the judge absolutely did not say that in any of the quotes in the article; in the hearing on both parties motions for partial sunmary judgement, he both said that would be the case if the plaintiffs proved certain facts and raised doubts that they have the evidence to prove them.

Re: Judge said Meta illegally used books to build its AI

#108

Earlier quoted context omitted.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

You can train yourself with every book at any library. You can also train yourself on a large number of movies and TV shows for a small monthly fee. Where your analogy goes wrong is you're saying you want to "[Circumvent] payment to obtain copyright material for training" to use Workaccount2's words.

So Meta borrowed every book from a library and paid to obtain all of the movies and TV shows? They kept only one copy of every book at any time on their system?

Because I'm certainly not allowed to photocopy a library book in its entirety. And I guarantee you a Netflix subscription doesn't allow me to keep a copy of a movie on my hard drive and use it for training man or machine.

Re: Judge said Meta illegally used books to build its AI

#109

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

> AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. IANAL, but it doesn't look that hard. On first glance this is a fair use issue. What an LLM spits out is pretty clearly transformative use. But the fact that it pulls not only the entirety of the work, but the entirety of MOST works means that the amount is way beyond what could be fair use. Plus it's commercial use.…

> the fact that it pulls not only the entirety of the work

What do you mean by "pulls"?

What matters in traditional fair use is how substantially your output copies the work (among other factors). Your input is generally assumed to be reading/watching/listening to the entire work, and there is no problem with that.

Post reply on HN