Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…
> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?
Judge said Meta illegally used books to build its AI
161–170 of 352 posts
Re: Judge said Meta illegally used books to build its AI
#162Earlier quoted context omitted.
> But also, you are allowed to make them. Not of physical media. You're allowed to make archival copies of digital media. > Or reading a book via a computer would be illegal No you purchased a license (or your library did, in the case of e-borrowing) to read the book on a computer. That makes it legal.
I am allowed to point a webcam at my physical book and read off the screen, even though that makes digital copies of all the text.
This scenario seems quite contrived but is there an actual court precedent allowing it? I'm 100% confident no one will ever prosecute you for doing it but that's not the same thing as "allowed".
In another thread I already posted about https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In....
This case was about pointing at and making a copy to transmit it to . The Supreme Court held it to be illegal. There was a lot of money on the line, which is why it went so far.
Re: Judge said Meta illegally used books to build its AI
#163Earlier quoted context omitted.
If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.
> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. End of the road for major AI companies, and hopefully something better can be created once it's declared illegal without any murky waters. There are LLMs trained on data that isn't illegally obtained, OLMo by Ai2 is one such model, that is actually open source and uses open…
Re: Judge said Meta illegally used books to build its AI
#164Re: Judge said Meta illegally used books to build its AI
#165Earlier quoted context omitted.
The FairTrained models claim to train with only public domain and legal works. Companies are also licensing works. This company has a lawful, foundation model: https://273ventures.com/kl3m-the-first-legal-large-language-... So, it's really the majority of companies breaking the law who will be affected. Companies using permissible and licensed works will be fine. The other companies would finally have to buy large co…
I don't know? Not really sure a claim is good enough. I don't know that you can just go into court and say, "Trust me, I don't use copyrighted material." And I also can't see any way, other than providing training data and training an identically structured model on that data, that a company can conclusively show that they got the weights in an allegedly copyright free model from the copyright free training data a co…
There is benefit to using them, though. For one, they've tried really hard to be legal. That sets a positive example, shows good faith if they were sued, and reduces risk for those using them (good faith on our part). Also, one can be sure that they can ditch or replace any outputs in the long term if they're ruled illegal. So, we try not to use the A.I.'s in a way where losing access to them seriously damages our business.
That's the best I can offer until legal reforms happen.
If training, one can train it in Singapore on material you he or she has legal access to. Their law pretty much let's you use anything for AI purposes so long as you legally can access it yourself. To further reduce the risk, they should crawl it themselves, too, taking care to avoid risky sources.
Re: Judge said Meta illegally used books to build its AI
#166Earlier quoted context omitted.
Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.
It's almost like information wants to be free
Re: Judge said Meta illegally used books to build its AI
#167Earlier quoted context omitted.
If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.
> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…
Please do!!
Re: Judge said Meta illegally used books to build its AI
#168Earlier quoted context omitted.
I am allowed to point a webcam at my physical book and read off the screen, even though that makes digital copies of all the text.
FWIW the essence of copyright law AIUI is: copying is not permitted, unless done in a form explicitly allowed by the license holder. This scenario seems quite contrived but is there an actual court precedent allowing it? I'm 100% confident no one will ever prosecute you for doing it but that's not the same thing as "allowed". In another thread I already posted about https://en.wikipedia.org/wiki/American_Broadcasting…
Your example involves transmission and mine doesn't, and that's a whole different can of worms.
Also the result of that case was self-contradicting so it's not a great basis to build too much logic upon.
Re: Judge said Meta illegally used books to build its AI
#169AI hucksters vs. the Copyright Cartel. When two evil villains fight, who do you root for? Here's hoping they somehow destroy each other.
Neither of them died, though, both parties just kept all the books from the public and used them for their own purposes, while normal people had to squirrel them away and trade them illegally. It's the Tech Cartels vs. the Copyright Trolls. It'll end up as a romance.
Re: Judge said Meta illegally used books to build its AI
#170I'm wondering if authors are making the same mistakes that the music industry did with Napster and kazaa. Using AI has led to more book purchases for me. If I discover and enjoy a book via AI I'm more inclined to buy it. The cats out of the bag, so pet him.
Can you share more about how you sample books with AI?