Live data from Hacker News

The Battle over Books3

wired.com

101–110 of 135 posts

Re: The Battle over Books3

#101
post #95
post #42

Earlier quoted context omitted.

I always took democratize access to X to mean “I would like to give many people the opportunity to give me lots of money by buying my product.”

I saw a funny tweet recently from a guy who was spending $187 a month to not learn art. He was subscribed to like 9 different ai subscription services

Unless he is in a third world country, the value of your time spent trying to 'learn art' would be much, much higher than that.

Re: The Battle over Books3

#103
Books3 is the easy/fast/cheap method but if the quality of the model really brings in some sort of revenue is there anything stopping a company from buying/checking out of the library all these books, scanning and ORC'ing them and adding to the model the hard way?

Re: The Battle over Books3

#106

Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…

I've developed a similar mistrust of the term. Blogged about it a while back https://writing.kemitchell.com/2023/01/05/Type-Error-Democra...

Re: The Battle over Books3

#107
>He sees the widespread practice of training AI on copyrighted data as outrageous, and finds it infuriating that this behavior gets defended with claims that it’s democratizing access to information. “Open source doesn’t mean you took a bunch of people’s shit and gave it away for free,” he says. “That's theft.”

>Whether the defendant had purchased a signed copy or flagrantly shoplifted a dog-eared paperback wouldn’t matter during arguments over whether The Bedwetter, Too was a derivative rip-off or a transformative parody.

This strikes at the heart of why this case is about to be laughed out of court.

The argument the plaintiffs are making is that ChatGPT is a "derivative work", i.e. letting people use the software is akin to distributing carbon copies of the book at issue, with at most slight modifications (typical derivative works include translations, screenplay adaptations, etc.).

Since ChatGPT obviously cannot literally produce the full text of the book on command, the very strained position they're trying to advance is that short, several-paragraph summaries constitute a derivative work.

That is to say, they're arguing that writing, say, a review, or a book report, is an act of copyright infringement tantamount to taking a book, translating it into Japanese, and selling that translation.

It's a deeply stupid and wrongheaded argument, and it deserves to die a quick death.

Re: The Battle over Books3

#108
post #55

Earlier quoted context omitted.

> And if those creators were instead to try to introduce legislation, the AI companies would risk losing access to content from small creators without the means to sue too. this is the point of a class action suit isn't it? if it turns out training isn't fair use then Microsoft/Google/OpenAI will suddenly have class action suits for billions if not trillions of damages against them ($150,000 damages per willful infri…

Have you never seen the outcome of a class action? They’re all slaps in the wrist, less than speeding tickets, and the action members get like a free hotdog or red bull as compensation if they’re lucky

The tobacco companies paid out $246B (That's b as in billions) in that class action.

Re: The Battle over Books3

#109
post #5

Does anyone have more content about what makes Books3 so special relative to Bibliotik? Was it processed somehow, or just compiled into a single file? I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that…

>It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name.

>Surely there must be more to the story?

I asked him a while ago about this. The story is as simple as it sounds:

https://news.ycombinator.com/item?id=36194455

Re: The Battle over Books3

#110
post #107

>He sees the widespread practice of training AI on copyrighted data as outrageous, and finds it infuriating that this behavior gets defended with claims that it’s democratizing access to information. “Open source doesn’t mean you took a bunch of people’s shit and gave it away for free,” he says. “That's theft.” >Whether the defendant had purchased a signed copy or flagrantly shoplifted a dog-eared paperback wouldn’t…

Writing a review or a book report is very much creating a derivative work in copyright law. Copyright law then says that these derivative works are a fair use. It does not follow that other derivative works that you personally feel are as serious are also fair use.
Post reply on HN