Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

111–120 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#111

Earlier quoted context omitted.

If you read a book and later understand its plot but can only explain it in your own words, did you copy it? The model isn’t storing the book.

I just happen do read the Phoenix Technologies wikipedia page a few days ago. This company is known for developing BIOS software for computers. Maybe you've seen their logo when you first turn on your computer. In early computing, everything was closed sourced. Quoting the wikiepdia page, To develop a legal BIOS, Phoenix used a clean room design. Engineers read the BIOS source listings in the IBM PC Technical Referen…

Perhaps Phoenix just looked at the potential adversary (IBM) and decided to approach the project in an exceedingly cautious way, knowing that IBM could litigate it forever if there were any plausible argument that they "copied" even a line of code.

Re: Judge said Meta illegally used books to build its AI

#112

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

The US Federal government operates with the rule that if human eyes don't look at it it doesn't count as a copy or looking at it. This allows them to unconstitutionally spy and log all people's telecommunications. Applying it here it seems pretty clear that corps are within the established bounds. As are any human persons that want to train an LLM this way.

Someone engaged in large-scale unconstitutional spying does not give two fs about incidentally doing some copyright violations to achieve the spying. These are entirely orthogonal considerations.

Re: Judge said Meta illegally used books to build its AI

#113
post #77

Earlier quoted context omitted.

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

You point to Chinese companies disregarding any rules if there is value to be had in AI, while in the US, AI companies going to get 500 billion investment and a whistleblower is dead.

US AI companies will either make sure that a similar ruling will never be made or they will ignore it and pay the fines. They won't let anybody stop the gravy train.

Re: Judge said Meta illegally used books to build its AI

#114
post #77

Earlier quoted context omitted.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

China found the perfect way to disrupt US tech, releasing open source versions of it for free or at least cheaper. Most of US tech is built on open source anyways and with the pace YC is investing in open source alternatives, it will win out in most niches. My fear is that the US tech won’t be able to compete with state sponsored open source out of China and will move to ban open source or suppress it somehow.

Also, the Chinese work is legit. DeepSeek introduced a whole bag of new techniques like GRPO, and released quite a bit of good open source tooling.

And Alibaba's Qwen team seems to be quite genuinely talented at "small" models, 32B parameters and below. Once you get Qwen3 properly configured, it punches well above its "weight class." I'm still running real benchmarks, but subjectively, it feels like the 32B model performs somewhere between 4o-mini and 4o on "objectively measureable" tasks. It's a little "stodgy" and formal by default, though. We'll see what it looks like when people start fine-tuning it.

If the US dropped off the planet, it would maybe set LLM technology back a year.

Re: Judge said Meta illegally used books to build its AI

#115
post #30

Earlier quoted context omitted.

> but may be entitled to use them under fair use. Why? Was it legal for me to download copyrighted songs from Limewire as "fair use" ? Because a few people were made examples of. I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;)

I don't believe anyone was ever penalized for downloading only uploading which seems like a pretty similar principle to what the judge is saying here.

Heh. People were penalized for merely creating search engines that happened to link to songs. Supposedly the RIAA accepted the offer of a 20-something’s life savings, but only if they switched their major from CS to something else. I believe it, having witnessed those times.

Re: Judge said Meta illegally used books to build its AI

#116

Earlier quoted context omitted.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

There’s something called a substantive transformation test in copyright law. When you write a summary of a book, you don’t infringe on copyright because it’s a “substantial transformation”. This goes along with the idea that you can copyright the text but not the ideas it expresses. When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong…

The counterargument to that is model training is impossible without making copies. That's not true for humans.

Re: Judge said Meta illegally used books to build its AI

#117

AI hucksters vs. the Copyright Cartel. When two evil villains fight, who do you root for? Here's hoping they somehow destroy each other.

I can only root for them both to lose.

Letting Meta launder copyrighted works to make billions, while threatening the rest of us over the most trivial derivative work, sounds like the worst outcome to me.

Copyright is a mistake. It demands that we compete instead of collaborate. LLMs don't provide enough utility to deserve special treatment in these circumstances. If anyone can infringe copyright, then everyone should be able to.

Re: Judge said Meta illegally used books to build its AI

#118

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

> Circumventing payment to obtain copyright material for training - Unambiguously illegal.

The judge in this case seems to disagree with you, not accepting the premise that downloading the material from pirate sites for this use inherently gets the plaintiffs an out from having to address fair use defense as to the actual use.

> the plaintiffs want to also tie in the former.

No, the defense wants to and the judge hasn't let the plaintiffs avoid it the way you argue they automatically can.

Re: Judge said Meta illegally used books to build its AI

#119

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another.

End of the road for major AI companies, and hopefully something better can be created once it's declared illegal without any murky waters.

There are LLMs trained on data that isn't illegally obtained, OLMo by Ai2 is one such model, that is actually open source and uses open data for training. Just because it's "very difficult" for OpenAI et al shouldn't be an argument to force them to behave ethically anyways. If they cannot survive acting legally, then so be it, sucks for them.

Re: Judge said Meta illegally used books to build its AI

#120

Earlier quoted context omitted.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.

Copyright is defined in law and as the original poster stated, whether this is 'copying' as defined by copyright law is legally ambiguous. Copyright doesn't protect against all forms of duplication. For instance, you own the copyright to your post and grant HN a license to offer copies of it. I have no direct license from you to copy the content of your post; but I can copy it to memory, copy a cache to disk, and mak…

> For instance, you own the copyright to your post and grant HN a license to offer copies of it.

It’s not a good example, because if you grant a license you give them the right to make copies. The problem is not when Meta got licenses, it’s when they did not.

Post reply on HN