Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

71–80 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#71
post #3

I think the headline is a bit misleading. Mets did pirate the works but may be entitled to use them under fair use. It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue…

The problem is that "harm" as defined by copyright law is strictly limited to loss of sales due to breach of that copyright; it makes no allowment (that I know of) to livelihoods lost by the theft of the work indefinitely, as AI boosters suggest their tools can do (replace people). The way this court case is going, it's an uphill battle for the plaintiffs to prove concrete harm in that very narrow context, when the r…

AI Boosters can suggest whatever nonsense strikes their fancy, and creatives can give into fear for no reason, but the best estimates we have from the BLS is that the there are careers and ongoing demand for artists, writers, photographers.

Regardless, deep learning models are valuable because they generalize within the training data to uncover patterns and features and relationships that are implicit, rather (simply) present with the data. While they can return things that happen to be within the training set, there is no reason to believe that any particular output is literally found there or is something that could be attributable, or that a human would ever attribute. Human artists also make meaning from the broad texture of their life experiences and general diffuse unattributable experience of culture.

Sure, this is something a random artist is unlikely to know, but if they are simply refusing to pick up a useful tools that can't give credit--say avoiding LLMs for brainstorming, or generative selection tools for visual editing, or whatever, their particular careers will be harmed by their incurious sentimentality, and other human artists will thrive because they know that tools are just tools, and it is the humans using the tools that make meaning that people care about.

Re: Judge said Meta illegally used books to build its AI

#72

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

I have a weird controversial view on this in terms of how to legally do it, and that is, for your 1 model, you should be only required to buy a digital copy of the work, maybe publishers should make digital copies that are tailored for LLMs to churn through, but then price it at a reasonable rate, and make the format basically perfect for LLMs.

Re: Judge said Meta illegally used books to build its AI

#73
post #67

Earlier quoted context omitted.

The FairTrained models claim to train with only public domain and legal works. Companies are also licensing works. This company has a lawful, foundation model: https://273ventures.com/kl3m-the-first-legal-large-language-... So, it's really the majority of companies breaking the law who will be affected. Companies using permissible and licensed works will be fine. The other companies would finally have to buy large co…

I don't know? Not really sure a claim is good enough. I don't know that you can just go into court and say, "Trust me, I don't use copyrighted material." And I also can't see any way, other than providing training data and training an identically structured model on that data, that a company can conclusively show that they got the weights in an allegedly copyright free model from the copyright free training data a co…

Civil courts work by you proving damages (at least in the USA), not by you going on fishing expeditions because they "might" have done something.

So good luck finding the thing that looks exactly like your copyrighted work that's not in the corpus, if you can yeah, you might be able to prove it.

At the end of the day its like a lot of business, where a liability shell game is played out, and if the chain of evidence cant be drawn quite brightly then lawsuits would be frivolous at best.

Re: Judge said Meta illegally used books to build its AI

#74

Earlier quoted context omitted.

> AI doesn't actually directly copy the material it trains on Of course it does. Large models are trained on gigantic clusters. How can you train without copying the material to machines in the cluster?

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Copyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training?

If we're going that way, let me torrent every movie and TV show ever to "train" myself.

Re: Judge said Meta illegally used books to build its AI

#75
post #67

Earlier quoted context omitted.

The FairTrained models claim to train with only public domain and legal works. Companies are also licensing works. This company has a lawful, foundation model: https://273ventures.com/kl3m-the-first-legal-large-language-... So, it's really the majority of companies breaking the law who will be affected. Companies using permissible and licensed works will be fine. The other companies would finally have to buy large co…

I don't know? Not really sure a claim is good enough. I don't know that you can just go into court and say, "Trust me, I don't use copyrighted material." And I also can't see any way, other than providing training data and training an identically structured model on that data, that a company can conclusively show that they got the weights in an allegedly copyright free model from the copyright free training data a co…

[deleted]

Re: Judge said Meta illegally used books to build its AI

#76
post #67

Earlier quoted context omitted.

The FairTrained models claim to train with only public domain and legal works. Companies are also licensing works. This company has a lawful, foundation model: https://273ventures.com/kl3m-the-first-legal-large-language-... So, it's really the majority of companies breaking the law who will be affected. Companies using permissible and licensed works will be fine. The other companies would finally have to buy large co…

I don't know? Not really sure a claim is good enough. I don't know that you can just go into court and say, "Trust me, I don't use copyrighted material." And I also can't see any way, other than providing training data and training an identically structured model on that data, that a company can conclusively show that they got the weights in an allegedly copyright free model from the copyright free training data a co…

I do hope people are still innocent until proven guilty?

If you did not use copyrighted materials for training, people will not be able to prove that you did, and that should be good enough.

Re: Judge said Meta illegally used books to build its AI

#77

Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…

If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another.

I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.)

If the US decides to unilaterally shut down LLMs, that just means that the rest of the world will route around us. Whether this is good or bad is another question.

Re: Judge said Meta illegally used books to build its AI

#78

Earlier quoted context omitted.

The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)

Copyright was invented (in its modern form) by corporations. It will be uninvented if need be for corporations.

The golden rule strikes again!

Re: Judge said Meta illegally used books to build its AI

#79
post #39
post #7

I'm wondering if authors are making the same mistakes that the music industry did with Napster and kazaa. Using AI has led to more book purchases for me. If I discover and enjoy a book via AI I'm more inclined to buy it. The cats out of the bag, so pet him.

Can you share more about how you sample books with AI?

It can tell you about authors, books, useful techniques, etc. If it cites references, that can generate page views on their site ir sales. It can also replace that, though, with AI supplier benefiting commercially.

Re: Judge said Meta illegally used books to build its AI

#80
post #30

Earlier quoted context omitted.

> but may be entitled to use them under fair use. Why? Was it legal for me to download copyrighted songs from Limewire as "fair use" ? Because a few people were made examples of. I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;)

I don't believe anyone was ever penalized for downloading only uploading which seems like a pretty similar principle to what the judge is saying here.

This is why Meta didn't seed.[0]

[0] https://torrentfreak.com/meta-says-it-made-sure-not-to-seed-...

Post reply on HN