Live data from Hacker News

Thomson Reuters wins first major AI copyright case in the US

wired.com

91–100 of 188 posts

Re: Thomson Reuters wins first major AI copyright case in the US

#91

Earlier quoted context omitted.

License what? Every available copyrighted work? Even getting a tiny fraction is not practical. To the contrary, this just means companies can't make money from these models. Those using models for research and personal use wouldn't be infringing under the fair use tests.

> License what? Every available copyrighted work? Even getting a tiny fraction is not practical. Maybe the strategy is something like this: 1) Survive long enough/get enough users that killing the generative AI industry is politically infeasible. 2) Negotiate a compromise similar to the compulsory mechanical royalty system used in the music business to “compensate” the rights holders whose content is used to train th…

> the basis that figuring out what IP went into what model output is just too hard, so instead they just agree to distribute it to whomever is on the New York Times best seller list at any given moment.

the long tail exists, and there will always be a threshold for payments due to rights holders.

it used to be (like 10 years ago so i might not remember the details exactly) that if you earned less than £1 from youtube performing music rights in a quarter then any money you earned was put back into the pot and redistributed to those earning over £1.

it just wasn’t worth the cost to keep track of £0.00001 earnings for all the rights holder in the bottom of the long tail each quarter, or to pay the bank fees when the eventually earn £0.01 that can be paid to them.

definitely not perfect, but at least some people were getting paid, instead of none.

also, youtube’s data they gave us was fairly shit (video title, url). so that didn’t help. nor did the lack of compute/data proc infrastructure/skills. was historically a manual spreadsheet job trying to work out who to cut.

i had to do it a few times :/

edit —

> The biggest AI companies could even run the enforcement cartels ala BMI/ASCAP to compute and collect royalties owed.

what could happen, for music at least, is the same thing that happened with youtube, mashed up with live music analogies.

a licensing negotiation with BMI/ASCAP/PRS, and maybe major publishers directly if they get frustrated with the PROs. then PROs will use sampling of other revenue streams to work out what the likely popular things are for AI. then divvy up whatever the lump sum is between the most popular songs.

we used to do this for live music. i had to generate the sampled dataset in microsoft access each year and weed out the all the radio stings.

sorry for costing you a million pounds that one year ed sheeran :/

Re: Thomson Reuters wins first major AI copyright case in the US

#92
post #62
post #33

See. The fair-use excuses that the AI proponents here were trying to hang on to for dear life have fallen flat on this ruling. This is going to be one of many cases in which there will be licensing deals being made out of this to stop AI grifters claiming 'fair use' to try to side-step copyright laws because they are using a gen AI system. OpenAI ended up paying up for the data with Shutterstock and other news source…

once the case law is set, I look forward to suing everyone that's ever trained a model for $300,000 PER WORK each time they ingested my code from GitHub whoever wrote those indemnity policies is going to regret it

> once the case law is set, I look forward to suing everyone that's ever trained a model for $300,000 PER WORK each time they ingested my code from GitHub

Didn't you already share it on GitHub royalty-free?

Re: Thomson Reuters wins first major AI copyright case in the US

#93

Earlier quoted context omitted.

This is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote…

> That, plus the fact that Ross was a directly competing product, is what I see as really driving this decision. The "competing product" thing is probably the most extreme part of this opinion. The most important fair use factor is if the use competes with the original work, but this is generally implied to be directly competes, i.e. if you translate someone else's book from English to French and want to sell the tra…

No that is not an extreme interpretation of the fair use factors. This is a routinely emphasized factor in fair use analyses for both copyright and trademark. School fair use is different because that defense is written into the statute directly in 17 U.S.C. § 107. Also, § 108 provides extensive protections for libraries and archives that go beyond fair use doctrines.

The idea that the schools are encouraging the students to compete with the original authors of works taught in the classroom is fanciful by the meaning that courts usually apply to competition. Your example is different from this case in which Ross wanted to compete in the same market against West offering a similar service at a lower price. Another reason that the schools get a carveout is because it would make most education impractical without each school obtaining special licenses for public performance for every work referenced in the classroom.

But maybe that also provokes the question as to if schools really deserve that kind of sweetheart treatment (a massive indirect subsidy), or does it over-privileges formal schools relative to the commons at large?

Re: Thomson Reuters wins first major AI copyright case in the US

#94
post #56

Earlier quoted context omitted.

If that mechanical process is not reversible, then it's not a copyright violation. For instance, I can compute the SHA256 hashes for every book in existence and distribute the resulting table of (ISBN, SHA256) and that is not a copyright violation.

That's actually within the other fair use factors. So your hash table is fair use because its transformative and doesn't substitute for the original work. I edited my post to make it a bit clearer.

It's actually even less than fair use, it's non-copyright use: one-way hashes are intentionally designed to eliminate the creative element and output random looking data.

Re: Thomson Reuters wins first major AI copyright case in the US

#95
post #39
post #23

Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model th…

If the copyright holders win, the model giants will just license. This effectively kills open source, which can't afford to license and won't be able to sublicense training data. This is very bad for democratized access to and development of AI. The giants will probably want this. The giants were already purchasing legacy media content enterprises (Amazon and MGM, etc.), so this will probably further consolidation an…

Open source models can crowdsource open source training data. This was done for RNNoise for example.

Re: Thomson Reuters wins first major AI copyright case in the US

#96
post #62

Earlier quoted context omitted.

once the case law is set, I look forward to suing everyone that's ever trained a model for $300,000 PER WORK each time they ingested my code from GitHub whoever wrote those indemnity policies is going to regret it

> once the case law is set, I look forward to suing everyone that's ever trained a model for $300,000 PER WORK each time they ingested my code from GitHub Didn't you already share it on GitHub royalty-free?

no, it was under a very specific license that required attribution

and other than that, All Rights Reserved

Re: Thomson Reuters wins first major AI copyright case in the US

#97
post #14

Earlier quoted context omitted.

> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substa…

As long as someone give me the software software to run my business, that person might be in violation of copyright but I'm in the clear. Simply running my business on illegally distributed copyrighted text/software/movie should not be copyright infringement.

If you buy a machine that prints copies of copyrighted books (built into the machine), and you use that machine and then distribute the resulting copies, and the machine didn't come with a license allowing you to do so, I'm pretty sure that you are liable as well.

At least some current AI providers, however, come with terms of service that promise that they will cover any such legal disputes for you.

Re: Thomson Reuters wins first major AI copyright case in the US

#98
post #16

Earlier quoted context omitted.

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

> A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. > I remember talking with friends 30 years ago I'd say you're pretty old. How many years of training did it take for you to start producing good output? The leason here is we're kind of meta-trained: our minds are primed to pick up new things quickly by abstracting them and relating them to things we alread…

That's the point I think. It should be possible to require orders of magnitude less data to create an intelligence, and we are far from achieving that (including achieving AGI in the first place even with those huge amounts of data).

Re: Thomson Reuters wins first major AI copyright case in the US

#99
post #67

Earlier quoted context omitted.

if you have an MNIST classifier that just takes in images, and spits out a probability of digits 1-9, that wouldn't necessarily be generative, if it is only capable of modeling P(which digit | all pixels). But many other types of model would give you a joint distribution P(which digit, all pixels), so would be generative. Even if you only used it for classification. https://en.wikipedia.org/wiki/Generative_model I gu…

You can derive the latter information (the joint distribution), given the former and a prior over "all pixels"-like data. So, the defining feature of "generative" models is that they feature a prior over their input data?

Generative models model the data, whether that is p(x) or p(x,y) or (x,y,z) etc.

Re: Thomson Reuters wins first major AI copyright case in the US

#100
post #16

Earlier quoted context omitted.

> the current method for training requires this volume of data This is one of those things that signal how dumb this technology still is - or maybe how smart humans are when compared to machines. A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. I remember talking with friends 30 years ago about how it was inevitable that the brain would eventually be fully…

> A human brain doesn't need anywhere close to this volume of data, in order to be able to produce good output. Maybe not directly, but consider that our brains are the product of million of years of evolution and aren't a blank slate when we're born. Even though babies can't speak a language at birth, they already have all the neural connections in place in order to acquire and manipulate language, and require just…

There can't be that much pre-built into the brain. There isn't that much dna, and only a portion of it can be going to the design of the brain.

A lot of what we're able to do has to be from some sort of generic capability.

Post reply on HN