Earlier quoted context omitted.
An LLM can be trained to find relevant knowledge online. It doesn’t have to be trained on the entirety of all existing knowledge.
> An LLM can be trained to find relevant knowledge online. Why do you think chatGPT lost its Web Search plugin lately? Copyright lawsuits. You can't even use copyrighted content in the prompt because it will make the model makers liable.
The Battle over Books3
131–135 of 135 posts
Re: The Battle over Books3
#132Earlier quoted context omitted.
No, as you can see from your very definition. But here's a good example: If you take a book and turn it into a movie, that's a derivative work. Anyone can see the direct resemblance -- the transformation or adaptation . But if you take a book, convert each letter to a number, add up the numbers that make each sentence, and then sell that as a list of "random" numbers, that's not a derivative work. The end result is s…
Weights are simply a lossy compression of the training data set. Now, I understand the argument that perhaps the specific work has been homeopathically diluted down to nothingness in the weights and so therefore has only been used to contextualise the compression process of other works, but if the weights can be reasonably used to generate copyright infringing text (and condensations and abridgements and transformati…
No they're not -- they're more like the dictionary generated to produce a lossless compressed data set. But then we throw out the compressed data itself, and keep only the dictionary.
> but if the weights can be reasonably used to generate copyright infringing text (and condensations and abridgements and transformations are explicitly listed in the law, verbatim copying is not necessary)
First of all, they haven't been shown to substantially generate infringing text that aren't the kinds of short snippets covered by fair use. And my previous comment already explained that longer texts are not going to happen, for both legal and economic reasons.
But secondly, you're wrong about "condensations and abridgements and transformations". You can absolutely sell a page-long summary of a book without getting permission, for instance. What do you think things like CliffsNotes are all about? Or all those two-page "executive summaries" of popular busines books?
You can't abridge a 1,000 page book to 500 pages and sell that, but you can summarize its ideas in a page and sell that. Which is basically the approximate level of understanding that LLM's seem to absorb.
Re: The Battle over Books3
#133Earlier quoted context omitted.
> Does anyone have more content about what makes Books3 so special relative to Bibliotik? They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.) > This is the kind of effort that could have been done anonymously Sure, its something each group training an…
So his contribution consisted of putting Bibliotik into plaintext format? Or is there more to it?
Re: The Battle over Books3
#134From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?
39.52Gb gzipped and 108.79GB uncompressed.
39,516,981,435 books3.tar.gz -- 36.5% compression ratio
108,371,325,720 tar -- uncompressed
Recompressed with 7-zip and xz:
25,221,357,605 b3.7z # with flags: -m0=ppmd (23.3%)
27,077,329,052 books3.tar.xz # with flags: -e9 (25.0%)
To see what other slower compressors could do, I checked results from a random 1,000 books. MCM would achieve about half the original tar.gz file size, but it's very slow.
104,043,048 q-x11.mcm - 19.0% -- from q.tar
106,181,190 q-m9.mcm
114,564,722 q.tar.bsc-m03-b1000000000
118,636,713 q.tar.bsc-m03
130,621,411 q.ppmd -- 23.9%
137,874,349 q-s29.tar.lzip
138,274,566 qultra.7z
141,815,136 q.lzma
142,626,724 q9.xz -- 26.1%
146,786,015 q.brotli
146,933,603 q.p7zip
146,933,619 q.7z
149,371,106 q9.bz2
198,213,051 q.gz -- gzip -9 -- 36.3%
198,213,193 q.zip
227,074,912 q.lz4
228,463,000 q.lzop
546,201,600 q.tar
Re: The Battle over Books3
#135Earlier quoted context omitted.
I am not a lawyer, but it seems right to me to say that the weights are a derivative work of the training set. > A “derivative work” is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.…
No, as you can see from your very definition. But here's a good example: If you take a book and turn it into a movie, that's a derivative work. Anyone can see the direct resemblance -- the transformation or adaptation . But if you take a book, convert each letter to a number, add up the numbers that make each sentence, and then sell that as a list of "random" numbers, that's not a derivative work. The end result is s…
the problem with this as an example is that copyright would not apply to this transformative work, not the original author's copyright nor your new authorship because this transformative work contains no creative human expression (unless the original book was designed to add up to some fortune cookie, of course, in which case you have not transformed it)
A nuttier, chewier example would be retelling a litigious story like Moana ("consider the copyright, across all these leaves... make way!"), from the pig's perspective or something, and seeing what would fly and what wouldn't.