Live data from Hacker News

The Battle over Books3

wired.com

121–130 of 135 posts

Re: The Battle over Books3

#121
post #9
post #3

Earlier quoted context omitted.

My guess is the big players hope is to steal an enough content and then build a self training LLM based off synthetic content (rehashed original works) before the theft part matters. Not sure how close they are to achieving but this seems to be a common SV gamble. Steal or do something shady, raise enough money / power so by the time your noticed, you have the money to win in the courts. You already see the propagand…

An LLM can be trained to find relevant knowledge online. It doesn’t have to be trained on the entirety of all existing knowledge.

> An LLM can be trained to find relevant knowledge online.

Why do you think chatGPT lost its Web Search plugin lately? Copyright lawsuits. You can't even use copyrighted content in the prompt because it will make the model makers liable.

Re: The Battle over Books3

#122
post #120

Books3 is the easy/fast/cheap method but if the quality of the model really brings in some sort of revenue is there anything stopping a company from buying/checking out of the library all these books, scanning and ORC'ing them and adding to the model the hard way?

Like these guys? https://books.google.com/

Google Books corpus sure comes in handy now in the LLM era.

Re: The Battle over Books3

#123

Earlier quoted context omitted.

I mean that's part of the conversation that needs to be had. I would argue libraries are an unadulterated good, but it is generally considered at best unethical and at worst illegal to re-use content that isn't your own, at least without a proper citation. Then there's also the issue with things like art, music, and code. Where does the line fall with scraping Github, Soundcloud, DeviantArt, or Instagram and using th…

> but it is generally considered at best unethical and at worst illegal to re-use content that isn't your own, at least without a proper citation. No it's not at all, except in extremely limited circumstances. When George Lucas made Star Wars, did he cite all the Westerns and space opera serials and movies that influenced him? When you give a presentation at work on why you should move to a sharded database, do you c…

But there is no bright line test for what is and isn't fair use.

Re: The Battle over Books3

#124
post #94

Earlier quoted context omitted.

The level of disrespect for artists I've seen here and on other forums with regards to this technology has been staggering. The entitlement for work you didn't do is so gross.

As a fanfiction writer, I find the entitlement of authors and corporations over work that I created gross. I transformed their work into something else, advertising the original work and making a bigger market for it, and in exchange we get C&Ds and copyright claims. Someone writes a reskin of The Odyssey and then a corporation claims every derivative for 75 years.

I have a lot more sympathy for you than the corporate interests that are currently trying to profit off of your work without your knowledge.

Re: The Battle over Books3

#125
post #119
post #63

Earlier quoted context omitted.

yes, the lawyers always clean up the aim would be to make training on copyrighted material legally toxic and render all existing datasets and trained weights unlawful the damages are simply a bonus

How do you prove a set of weights was trained on, in part, any dataset? "We bought these weights in good faith from Digitus Tertius corp of bermuda."

trap phrases deal with that case very easily

Re: The Battle over Books3

#126

Earlier quoted context omitted.

I mean that's part of the conversation that needs to be had. I would argue libraries are an unadulterated good, but it is generally considered at best unethical and at worst illegal to re-use content that isn't your own, at least without a proper citation. Then there's also the issue with things like art, music, and code. Where does the line fall with scraping Github, Soundcloud, DeviantArt, or Instagram and using th…

> but it is generally considered at best unethical and at worst illegal to re-use content that isn't your own, at least without a proper citation. No it's not at all, except in extremely limited circumstances. When George Lucas made Star Wars, did he cite all the Westerns and space opera serials and movies that influenced him? When you give a presentation at work on why you should move to a sharded database, do you c…

I am not a lawyer, but it seems right to me to say that the weights are a derivative work of the training set.

> A “derivative work” is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted. A work consisting of editorial revisions, annotations, elaborations, or other modifications, which, as a whole, represent an original work of authorship, is a “derivative work”.

As I understand it, derivative works must be created with the legal use of the original work, or be fair use, otherwise they are infringing.

Re: The Battle over Books3

#127
post #65

Earlier quoted context omitted.

Most all humans communicate for a living. I don't see your point.

But not all and not most produce intellectual content for a living. And those who do you seem to be ok with ditching because if I read you correctly it's a small loss.

That's not what I was trying to convey, so I'll try again.

Humans generate a massive amount of natural language, and 99% of it never gets someplace where GPT/LLM training can consume it. If we can capture just a couple percent of that, then there will be no need for GPT/LLM to make use of content from those who don't want their writing to be consumed.

Re: The Battle over Books3

#128

Earlier quoted context omitted.

> but it is generally considered at best unethical and at worst illegal to re-use content that isn't your own, at least without a proper citation. No it's not at all, except in extremely limited circumstances. When George Lucas made Star Wars, did he cite all the Westerns and space opera serials and movies that influenced him? When you give a presentation at work on why you should move to a sharded database, do you c…

I am not a lawyer, but it seems right to me to say that the weights are a derivative work of the training set. > A “derivative work” is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.…

No, as you can see from your very definition. But here's a good example:

If you take a book and turn it into a movie, that's a derivative work. Anyone can see the direct resemblance -- the transformation or adaptation.

But if you take a book, convert each letter to a number, add up the numbers that make each sentence, and then sell that as a list of "random" numbers, that's not a derivative work. The end result is sufficiently transformed that copyright no longer applies. Ownership of the original work has no relevance.

And AI weights are like that. They're a complete transformation. They're not a derivate work. The only thing you have to make sure of is that they haven't been overtrained to the extent that they can regurgitate whole chapters of the texts they were trained on, for example. But that's not something they're currently able to do, and obviously copyright law will force companies to ensure it stays that way. (Not to mention that companies would do it anyways, due to the economic motivation of reducing model sizes to cut costs.)

Re: The Battle over Books3

#129

Earlier quoted context omitted.

I am not a lawyer, but it seems right to me to say that the weights are a derivative work of the training set. > A “derivative work” is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.…

No, as you can see from your very definition. But here's a good example: If you take a book and turn it into a movie, that's a derivative work. Anyone can see the direct resemblance -- the transformation or adaptation . But if you take a book, convert each letter to a number, add up the numbers that make each sentence, and then sell that as a list of "random" numbers, that's not a derivative work. The end result is s…

[deleted]

Re: The Battle over Books3

#130

Earlier quoted context omitted.

I am not a lawyer, but it seems right to me to say that the weights are a derivative work of the training set. > A “derivative work” is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.…

No, as you can see from your very definition. But here's a good example: If you take a book and turn it into a movie, that's a derivative work. Anyone can see the direct resemblance -- the transformation or adaptation . But if you take a book, convert each letter to a number, add up the numbers that make each sentence, and then sell that as a list of "random" numbers, that's not a derivative work. The end result is s…

Weights are simply a lossy compression of the training data set.

Now, I understand the argument that perhaps the specific work has been homeopathically diluted down to nothingness in the weights and so therefore has only been used to contextualise the compression process of other works, but if the weights can be reasonably used to generate copyright infringing text (and condensations and abridgements and transformations are explicitly listed in the law, verbatim copying is not necessary), or even answer substantial questions about it, then that shows that the weights included that data.

If I take a sound file and compress it down so it's poor quality but I can still make out the tune, that doesn't mean that I've avoided copyright law.

Post reply on HN