Live data from Hacker News

The Battle over Books3

wired.com

11–20 of 135 posts

Re: The Battle over Books3

#11
post #5

Does anyone have more content about what makes Books3 so special relative to Bibliotik? Was it processed somehow, or just compiled into a single file? I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that…

> Does anyone have more content about what makes Books3 so special relative to Bibliotik?

They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.)

> This is the kind of effort that could have been done anonymously

Sure, its something each group training an AI could do independently at greater aggregate cost until someone succeeds in taking the original source down, but not only would that be costlier, but it in would involve less transparency and comparability across model architecture, or at least required the transparent, comparable trained version to be different from the full version.

Re: The Battle over Books3

#13
post #5

Does anyone have more content about what makes Books3 so special relative to Bibliotik? Was it processed somehow, or just compiled into a single file? I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that…

> Does anyone have more content about what makes Books3 so special relative to Bibliotik? They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.) > This is the kind of effort that could have been done anonymously Sure, its something each group training an…

> Sure, its something each group training an AI could do independently at greater aggregate cost until someone succeeds in taking the original source down, but not only would that be costlier, but it in would involve less transparency and comparability across model architecture, or at least required the transparent, comparable trained version to be different from the full version.

I meant he could have uploaded it under a pseudonym rather than broadcasting to the world that he was the one doing the uploading.

Re: The Battle over Books3

#14

I lean to the side of AI progress over copyright, though I fear AI becoming stupid when incentive for original works maybe goes down, so I think society needs to figure out some social contract at least, like subsidies to keep new content flowing. For instance blogs will stop posting if chatbots in search grab all the answers directly from their sites bypassing all ad revenue.

> incentive for original works maybe goes down

I don't know why you would have that concern. LOTS of people write because they just want to share. Now broaden it out to include speech. LOTS of people talk - it's what we do.

Progress in AI will happen because humans like to express themselves. The challenge isn't copyright. It's figuring out how to capture the vast content that just isn't getting captured. Also, this is really only an issue for "new intelligence" - if you really think there is such a thing. Personally, I do not. I think like 99% of all human intelligence is in the out of copyright corpus.

Re: The Battle over Books3

#15
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

Is all the data on the internet from 2010 to 2025 that much more valuable than all the data on the internet from 2010 to 2022? Better data yes, but more data up to the present day? Do we really need that to keep improving AI?

GPT is already capable of incredible generalized language understanding and I'd wager we've long since hit diminishing returns from raw internet data. RLHF, fine tuning, and better (and more data efficient) architectures are what we need now.

Re: The Battle over Books3

#16
post #8

I wonder how this affects things like the Oxford English Dictionary, wiktionary, and other corpus-based linguistics, which rely on sample sentences of the word usage in order to determine the context. Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "ensh…

It shouldn't affect corpus linguistics but the other way around, because all these things have long ago (even before computers) been legally contested by authors and publishers w.r.t. what can be done to text by dictionary makers and corpora managers without the authors' permission, so unless new laws get passed, the current precedents establishing what's permissible for corpus linguistics would still be valid and also be relevant for treatment of machine learning models.

In essence, the long established principles for text analysis is that facts about text (concordances, collocation statistics, n-gram counts) are neither copyrightable nor derived work, and thus can be calculated, gathered, used and distributed even if copyright holders of the source data object. Now a court might judge that training a large language model is substantially different or that it's effectively the same, but such a decision wouldn't affect corpus linguistics and how they use sample sentences, only whether LLMs get the same treatment or not.

Re: The Battle over Books3

#17
From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?

Re: The Battle over Books3

#18

From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?

800gb

Re: The Battle over Books3

#19
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

Is all the data on the internet from 2010 to 2025 that much more valuable than all the data on the internet from 2010 to 2022? Better data yes, but more data up to the present day? Do we really need that to keep improving AI? GPT is already capable of incredible generalized language understanding and I'd wager we've long since hit diminishing returns from raw internet data. RLHF, fine tuning, and better (and more dat…

People have somehow conflated artificial intelligence with "oracle that knows everything" and so keeping models up to date with recent knowledge has become a must. Of course, that could be done by fine tuning techniques and better architectures that can outsource knowledge retrieval to tools, all very interesting ongoing work on that topic, but even these approaches require to not be "blocked" in their data retrieval tasks for them to work well.

I'm sure you can do lots of interesting research using outdated datasets but for companies creating products this will not be sufficient.

Re: The Battle over Books3

#20
post #9
post #3

Earlier quoted context omitted.

My guess is the big players hope is to steal an enough content and then build a self training LLM based off synthetic content (rehashed original works) before the theft part matters. Not sure how close they are to achieving but this seems to be a common SV gamble. Steal or do something shady, raise enough money / power so by the time your noticed, you have the money to win in the courts. You already see the propagand…

An LLM can be trained to find relevant knowledge online. It doesn’t have to be trained on the entirety of all existing knowledge.

Good luck when everything is paywalled
Post reply on HN