Earlier quoted context omitted.
It's also true for humans, you memorize only parts of what you read and see but you still had to view the whole thing first. The computer model is working differently of course but functionally it's the same idea.
God I hate this conversation so much. These cases have nothing to do with how the brain works.
Judge said Meta illegally used books to build its AI
311–320 of 352 posts
Re: Judge said Meta illegally used books to build its AI
#312Earlier quoted context omitted.
If you do that, it won't be able to give you a summary detailed enough to infringe anything.
It may give me a summary good enough that I don't have to buy the book, since it read the book. If there are any parts that aren't detailed enough for me, I can ask them to be expanded. If you're telling me that's not "infringing," you should follow what up with the argument for why it is not.
There are a lot of Wikipedia articles with summaries good enough "to not have to buy the book", because they were written by people who read the book. Should those be taken down?
Re: Judge said Meta illegally used books to build its AI
#313Earlier quoted context omitted.
I'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did no…
The source being illegal doesn’t make your use legal. Infact one could argue that it’s equally illegal or worse since a corporation knowingly engaged in illegal activity.
Just by being involved as a party does not make you culpable. Murderers are criminals, the murdered, less so.
Choosing to be a party might not make you culpable. You may be an active participant but unaware of the law breaking (being defrauded). Or the law may explicitly state that you can engage with people committing criminal acts and reap the benefits so long as you don't break those laws (or encourage them to be broken) yourself. Some forms of journalism are protected in this way.
Ultimately to have a case you have to state.
1. What law was broken 2. How an action by a party is in violation of that law. 3. That the action actually happened.
The largest problem with this case is not that 3. is in doubt but showing which 1. and 2. they are talking about.
Re: Judge said Meta illegally used books to build its AI
#314The title for this submission is somewhat misleading. The judge didn't make any sort of ruling, this is just reporting on a pretrial hearing. He also doesn't seem convinced as to how relevant downloading books from LibGen is to the case: > At times, it sounded like the case was the authors’ to lose, with [Judge] Chhabria noting that Meta was “destined to fail” if the plaintiffs could prove that Meta’s tools created s…
Better to read the submission before drawing conclusions rather than only the HN title. In this case the HN title has been editorialised.
The actual title of the article is "A Judge Says Meta's AI Copyright Case Is About `the Next Taylor Swift'"
"The judge didn't make any sort of ruling, this is just reporting on a pretrial hearing."
The HN title doesn't mention anything about a "ruling". Nor does the title chosen by Wired.
The subheading in the article reads "Meta's contentious AI copyright battle is heating up-and the court may be close to a ruling."
That is accurate. The Court will soon decide the SJ motions.
Reading the article leaves no chance of being mislead by any title:
"If Chhabria grants either motion, he'll issue a ruling before the case goes to trial-and likely set an important precedent shaping how courts deal with generative AI copyright cases moving forward."
Re: Judge said Meta illegally used books to build its AI
#315Let me make a clarifying statement since people confuse (purposely or just out of ignorance) what violating copyright for AI training can refer to: 1. Training AI on freely available copyright - Ambiguous legality, not really tested in court. AI doesn't actually directly copy the material it trains on, so it's not easy to make this ruling. 2. Circumventing payment to obtain copyright material for training - Unambiguo…
If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. What is inspiration? What is imitation? What is plagiarism? The lines aren't clearly drawn for humans... much less for LLMs.
Re: Judge said Meta illegally used books to build its AI
#316Earlier quoted context omitted.
> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…
China found the perfect way to disrupt US tech, releasing open source versions of it for free or at least cheaper. Most of US tech is built on open source anyways and with the pace YC is investing in open source alternatives, it will win out in most niches. My fear is that the US tech won’t be able to compete with state sponsored open source out of China and will move to ban open source or suppress it somehow.
Meta did this first I believe?
Re: Judge said Meta illegally used books to build its AI
#317And this is how Chinese model will win in long term, perhaps... They will be trained on everything and anything without consequences and we will all use it because these models are smarter (except for area like Chinese history and geography). I don't have the right answer on what can be done here to protect copyright or rather contributing back to authors of a paper without all these millions dollar wasted in lawsuit…
Re: Judge said Meta illegally used books to build its AI
#318I'm wondering if authors are making the same mistakes that the music industry did with Napster and kazaa. Using AI has led to more book purchases for me. If I discover and enjoy a book via AI I'm more inclined to buy it. The cats out of the bag, so pet him.
Re: Judge said Meta illegally used books to build its AI
#319Earlier quoted context omitted.
The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. Now that big capital wants to steal from individuals, big capital wins again. (Unrelatedly, has Boies ever won a high profile lawsuit? I remember him from the Bush/Gore recount issue, where he represented the Democrats.)
> The RIAA lawyers never had to demonstrate that copying a DVD cratered the sales of their clients. They just got high penalties for infringers almost by default. The argument for 'fair use' in DVD copying/sharing is much weaker since the thing being shared in that case is a verbatim, digital copy of the work. 'Format shifting' is a tenuous argument, and it's pretty easily limited to making (and not distributing) per…
This doesn't actually matter though does it? They still had access to a copy of the data in the first place to train the AI on
Since they likely did not pay a license to have access to the books they trained the AI on, then they violated copyright
The same way it would be violating copyright for a university student to pirate a textbook and learn from it
Re: Judge said Meta illegally used books to build its AI
#320Earlier quoted context omitted.
Derivative works are not generally allowed in many jurisdictions. Try releasing a cover song without clearing it first etc. Even using a recognizable sample will bite you Derivative works are tolerated in some cases like some manga or fanfics but it is a gray area and whenever the author or publisher wants to pursue it it is their full right to do it. Many do pursue it (You can get inspired by something, and this is…
> Try releasing a cover song without clearing it first etc. Even using a recognizable sample will bite you So… it’s complicated. This is one of the weird areas where music copyright and other copyright seem to differ in the US. In the US the situation is complex and there are a lot of weird special interests [0], but generally a composer/author of a song has the right to decide who first records and releases the song…
Which is compulsory for the performer too.
A derivative work like cover is sort of acceptable when it's performed by a person live for some audience (grey area but twitch sort of allows it. with a bunch of rules). As soon as you want to publish it you MUST have a license. And chatbot is a derivative work totally not performed live by a person for some audience
I saw great tracks that got taken down from all legal channels because they featured a sample from another song. Sometimes they remained up but mostly they were taken down. It is fully original publisher's discretion...