Live data from Hacker News

Judge said Meta illegally used books to build its AI

wired.com

281–290 of 352 posts

Re: Judge said Meta illegally used books to build its AI

#281
post #255
post #77

Earlier quoted context omitted.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

> If the US decides to unilaterally shut down LLMs, that just means that the rest of the world will route around us. You're talking as if they are some kind of nationalized or publically-owned asset, as opposed to a bunch of for-profit, privately-owned silos.

These cases can set precedents that basically shut down all all future "useful" AI implementations if the judges go too far. You can bet that CCP doesn't care one whit about US copyright law if leads to them leapfrogging us, it will definitely be j-johna-jameson-laugh.gif . I think that's the point of the commenter from 20000ft

Re: Judge said Meta illegally used books to build its AI

#282
post #55

Earlier quoted context omitted.

I'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did no…

Blizzard managed to get a copyright infringement win against a defendant company that merely accessed their game client (IP) in memory: a cheat reading values of player position IP that had been previously loaded by Blizzard itself https://en.wikipedia.org/wiki/MDY_Industries,_LLC_v._Blizzar... .

That link says that 'win' was a summary judgement that was reversed upon appeal.

It seems the case is ongoing, but the case against the company is not one of them committing copyright infringement, but that copyright infringement occurred that they encouraged, enabled, and profited from.

Re: Judge said Meta illegally used books to build its AI

#283
post #55

Earlier quoted context omitted.

I'm not sure if Meta did anything illegal in 2. either. I thought the copyright infringement was by the people who provided the copyrighted material when they did not have the rights to do so. I may be wrong on this, but it would seem a reasonable protection for consumers in general. Meta is hardly an average consumer, but I doubt that matters in the case of the law. Having grounds to suspect that the provider did no…

Even if your belief that only the person *providing* the content is liable, do you honestly think a single person found all the content, downloaded it, directly trained the model themselves, and then deleted the content? If at any step the content was given or shared to anyone else for any reason, have they not converted into a provider themselves?

That in itself is a complicated issue, but I would suspect that this does not count as copyright infringement. The entity in possession of the data does not change, the copies and manipulation are performed by employees but at no time do they own what they are manipulating.

If it were true that this constitutes a transfer of possession, then the targets for the lawsuit should be the individual employees, and I don't think anyone wants that.

Re: Judge said Meta illegally used books to build its AI

#284
post #58
post #54

Earlier quoted context omitted.

I'm curious what you mean by "in it's modern form". You seem to suggest there was a previous form that was not invented by corporations, but I don't believe that is the case.

Copyright was first established by governments and the Church prior to the invention of corporations.

> The concept of copyright first developed in England. In reaction to the printing of "scandalous books and pamphlets", the English Parliament passed the Licensing of the Press Act 1662, which required all intended publications to be registered with the government-approved Stationers' Company, giving the Stationers the right to regulate what material could be printed.

Source: https://en.wikipedia.org/wiki/Copyright

Okay, not a corporation, but a company.

Re: Judge said Meta illegally used books to build its AI

#285
post #77

Earlier quoted context omitted.

> If the former ever gets tested in court, it's the end of the road. All major AI companies have trained on copyrighted work, one way or another. I can absolutely guarantee you that neither DeepSeek nor Alibaba's highly talented Qwen group will care even a little bit, in the long run. Not if there's value to be had in AI. (And I can tell you down to the dollar what LLMs can save in certain business use cases.) If the…

> And I can tell you down to the dollar what LLMs can save in certain business use cases.) Please do!!

$3.50.

But seriously, LLMs are useless for anything except the most basic secretarial tasks and even there they still require as much reviewing as a high school intern. The reason executives want to use AI instead of humans is because AI is a capital expense that does not hit the financial the same way that wages do. (Capital expenses are excluded from EBITDA, which is the most common metric for measuring a company's financial performance. But it's become so popular for companies to push expenses to ITDA whenever they can that financial analysts are starting to push back and include ITDA in their analyses.)

In a nutshell, there are 3 ways to look at company financials: the PR way (EBITDA), the financial reporting way (GAAP, which includes ITDA), and the tax way (starts from GAAP or IFRS but with numerous rules on what items are included or excluded).

Re: Judge said Meta illegally used books to build its AI

#286
post #54

Earlier quoted context omitted.

I'm curious what you mean by "in it's modern form". You seem to suggest there was a previous form that was not invented by corporations, but I don't believe that is the case.

There were two, major differences in prior forms of copyright: 1. It protected works to reward authors during their lifetime. This was changed to lasting a long time after the author was dead. Then, also for corporations that were only persons on paper and theoretically immortal. This shift let companies squeeze money out of monopolized ideas for over a century rather than supporting artists and their creations. Inst…

> It protected works to reward authors during their lifetime.

This was the pitch by the (printing) companies that established copyright to gain a legal monopoly on copying.

> Instead of supporting the small fish, copyright law can reinforce the dominance of the sharks and whales.

This was always the case. It's just worse now.

Personally I'd love to see sane copyright laws. That is, abolishing copyright. Copyright only helps people in power stay in power. Oh, does Disney argue that you used a character of theirs in your story? Does it even matter if you did? You think you can take on Disney's lawyers? It doesn't matter if copyright is 5 years or 500. You're going to lose.

Art, and things erroneously treated like art under copyright like software, would be so much better if people could do it without fear of being a victim of copyright. Imagine if anyone could add their own flare to any story ever. Incredible.

Re: Judge said Meta illegally used books to build its AI

#287

Earlier quoted context omitted.

I would like to expand on this, since it seems to be a common misunderstanding. Lets imagine a hypothetical situation where one friend loans a book to another, who then makes a copy of it. The lender owns the book, and it is within his rights to loan it to whoever he wants. That is legal. Making this illegal would end libraries. The borrower is well within his rights to accept the book, and as the current owner he is…

Do you have any case law (other than Tivo or VHS time-shifting) that relates directly to books?

There was a relatively famous Google case regarding their digitization of books without the authors consent in 2015. Although it's not a perfectly analogous to this situation.

In Googles case they were digitizing the books (that they did not own), and publishing snippets for search users to help them find books and other material that weren't indexed on the web. The court found they had that right, but did place some pretty strict limits on them.

Still, Google was allowed to keep their database of scanned material despite not owning the originals.

Link: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Re: Judge said Meta illegally used books to build its AI

#288

Earlier quoted context omitted.

“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.

Why does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG b…

The keyword you’re looking for is “transformative”

Re: Judge said Meta illegally used books to build its AI

#289

Earlier quoted context omitted.

I have a weird controversial view on this in terms of how to legally do it, and that is, for your 1 model, you should be only required to buy a digital copy of the work, maybe publishers should make digital copies that are tailored for LLMs to churn through, but then price it at a reasonable rate, and make the format basically perfect for LLMs.

This is actually clever, let the market decide the price and the worth of each book for training. Pricing per model might be tricky, instead annual licensing for training might be better pricing structure. Very quickly all big publishers and big labs might find very precisely what the fair price is to pay per book/catalogue.

Yeah, you dont want to price them way too high to where nobody will pay for them, maybe even have the dreaded "Contact us for Pricing" thing setup.

Re: Judge said Meta illegally used books to build its AI

#290

Earlier quoted context omitted.

Copyright law does not restrict storing copyright information. It restricts distribution of copyright data without permission. So a computer can store and analyze data but cannot spit it out verbatim. If it spits it out under fair use clause, then it becomes debatable whether the new work is fair use.

Then why folks were arrested for filming in the cinemas? I don't think that's how the law works [1]: > 106. Exclusive rights in copyrighted works > Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following: > (1) to reproduce the copyrighted work in copies or phonorecords; And later: > 501. Infringement of copyright > (a) Anyone w…

Verbatim is not the point. “Transformative” is. You added nothing to the work when you film a movie at the theater. Google Books (Authors Guild v. Google, 2015) literally shows you photocopies of the pages but is considered transformative and fair use because they offer an easily searchable digital database (and also not harming the book selling market).
Post reply on HN