Live data from Hacker News

The Battle over Books3

wired.com

51–60 of 135 posts

Re: The Battle over Books3

#51

Interesting. When it's copyrighted works in a digital form (plain text, ePub, whatever) it's a legal issue. ROT13 the text is the data still a copyright violation? I imagine so since it's a trivial thing to restore it to a legally volatile form. I understand that an unresolved issue is whether, once ingested into an LLM, the trained LLM is in violation of copyright. One wonders if human readers too are in violation o…

> Is a "brain transplant" from one LLM to another a thing?

At least it's a thing for vision models and was also used to train OpenAI Five (the dota AI)[0]. So it probably applies to LLMs too.

[0] https://cdn.openai.com/dota-2.pdf (page 2)

Re: The Battle over Books3

#52

Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…

I think this is an oversimplification. The concern about entrenching the positions of big tech companies is, I think, genuine, and I do believe it's important to find ways to foster competition and opportunity for smaller players. Possibly the law needs to evolve and/or there need to be licensing solutions for this content that work both for creators, and for those looking to train models (and, in some sense, licensing payments that are "means-tested").

Re: The Battle over Books3

#53
post #8

I wonder how this affects things like the Oxford English Dictionary, wiktionary, and other corpus-based linguistics, which rely on sample sentences of the word usage in order to determine the context. Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "ensh…

Sample sentences like those in the OED are pretty much the textbook example of fair use.

They are obviously transformative (the whole "dictionary" part), copy factual information (use of a word, not what the sentence itself is saying), are not substantial (one sentence out of many thousands), and do not impact the original work's value (nobody would buy a dictionary instead of a novel because it contains a sample phrase from that novel). The OED cites its sources, which also strengthens its case.

Compare that to AI, which is more than happy to write a short story to the prompt "Write a story about Bucky and Captain America falling in love, and living happily ever after in a mountain cabin." (Transformative? Maybe. Factual? No. Substantial? Yes. Impacts value? Yes.) Works like the AI's output have been dealt with in lawsuits like Salinger v. Colting, and it simply is not allowed. The big question right now is: what about the AI model itself?

Re: The Battle over Books3

#54

Earlier quoted context omitted.

> Does anyone have more content about what makes Books3 so special relative to Bibliotik? They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.) > This is the kind of effort that could have been done anonymously Sure, its something each group training an…

So his contribution consisted of putting Bibliotik into plaintext format? Or is there more to it?

How dare you profane the contributions of a literal god of AI!

Re: The Battle over Books3

#55
post #34

What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators. Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legisla…

> And if those creators were instead to try to introduce legislation, the AI companies would risk losing access to content from small creators without the means to sue too.

this is the point of a class action suit isn't it?

if it turns out training isn't fair use then Microsoft/Google/OpenAI will suddenly have class action suits for billions if not trillions of damages against them

($150,000 damages per willful infringement, after all)

Re: The Battle over Books3

#56
post #50
post #49

Earlier quoted context omitted.

[flagged]

How so ? Copyright was introduced along with the print press to protect the works of authors. When you buy a copy of a book it's not exactly private property of the author that you buy, is it ?

Texts were copied even before the printing press.

See https://en.m.wikipedia.org/wiki/History_of_copyright for more information.

Re: The Battle over Books3

#57

I lean to the side of AI progress over copyright, though I fear AI becoming stupid when incentive for original works maybe goes down, so I think society needs to figure out some social contract at least, like subsidies to keep new content flowing. For instance blogs will stop posting if chatbots in search grab all the answers directly from their sites bypassing all ad revenue.

> incentive for original works maybe goes down I don't know why you would have that concern. LOTS of people write because they just want to share. Now broaden it out to include speech. LOTS of people talk - it's what we do. Progress in AI will happen because humans like to express themselves. The challenge isn't copyright. It's figuring out how to capture the vast content that just isn't getting captured. Also, this…

> LOTS of people write because they just want to share.

You just gave up the "for a living" group, who arguably produce overall better content (of course there are exceptions), and focused on hobbyists. I'd call that a self-defeat.

Re: The Battle over Books3

#58
post #55
post #34

What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators. Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legisla…

> And if those creators were instead to try to introduce legislation, the AI companies would risk losing access to content from small creators without the means to sue too. this is the point of a class action suit isn't it? if it turns out training isn't fair use then Microsoft/Google/OpenAI will suddenly have class action suits for billions if not trillions of damages against them ($150,000 damages per willful infri…

Have you never seen the outcome of a class action? They’re all slaps in the wrist, less than speeding tickets, and the action members get like a free hotdog or red bull as compensation if they’re lucky

Re: The Battle over Books3

#59
post #56
post #50

Earlier quoted context omitted.

How so ? Copyright was introduced along with the print press to protect the works of authors. When you buy a copy of a book it's not exactly private property of the author that you buy, is it ?

Texts were copied even before the printing press. See https://en.m.wikipedia.org/wiki/History_of_copyright for more information.

They were, usually to preserve the work, not distribute it on a large scale because there was no large scale. What's your argument ?

Re: The Battle over Books3

#60

Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…

Any time I see the phrase "democratize access," my spidey-sense starts tingling.

Funny, I get the same tingling feeling when I see the phrase "restrict access." I guess it's an unreliable signal, huh.

Post reply on HN