Live data from Hacker News

The Battle over Books3

wired.com

41–50 of 135 posts

Re: The Battle over Books3

#41

Earlier quoted context omitted.

books3.tar.gz itself is ~37gb compressed. Often really the entire "The Pile" dataset (composed of both the mostly compressed archives, along with a ~450gb compressed jsonl compilation of the data) is being discussed. That's around 825gb.

Out of curiosity why is jsonl so popular in the ML space?

Not speaking for ML, but I love jsonl as a distribution format. Lots of storage overhead vs a more appropriate bulk container, but it makes it trivial to stream, sample, concatenate, etc. Anything else (eg csv or parquet) is going to require better tooling and/or just a little shell magic to handle headers.

Re: The Battle over Books3

#42

Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…

I always took democratize access to X to mean “I would like to give many people the opportunity to give me lots of money by buying my product.”

Re: The Battle over Books3

#43

Earlier quoted context omitted.

> incentive for original works maybe goes down I don't know why you would have that concern. LOTS of people write because they just want to share. Now broaden it out to include speech. LOTS of people talk - it's what we do. Progress in AI will happen because humans like to express themselves. The challenge isn't copyright. It's figuring out how to capture the vast content that just isn't getting captured. Also, this…

> LOTS of people write because they just want to share. The blogging scene from the early millennium is now a shadow of its former self, and one of the most often stated reasons for abandoning blogging is “my site just wasn’t getting many views any more”. In a world where AI generated content abounds, there will be even fewer eyeballs on whatever one shares and therefore less feeling of reward for sharing. Moreover,…

Yah, blogs are already a rounding error in the corpus, and that has nothing to do with llms. Those of us who are still blogging are already doing it in spite of ~waves hand broadly at the world~.

I'm not truly sure that llms mean less eyeballs, though. They produce mediocre content in an arena where high quality content matters. There's already a massive pile of crap on the internet; it's already all about surfacing the relevant and the interesting bits.

Re: The Battle over Books3

#44

From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?

books3.tar.gz itself is ~37gb compressed. Often really the entire "The Pile" dataset (composed of both the mostly compressed archives, along with a ~450gb compressed jsonl compilation of the data) is being discussed. That's around 825gb.

That is shockingly approachable for a large fraction of English literature.

Re: The Battle over Books3

#45
post #5

Does anyone have more content about what makes Books3 so special relative to Bibliotik? Was it processed somehow, or just compiled into a single file? I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that…

> Does anyone have more content about what makes Books3 so special relative to Bibliotik? They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.) > This is the kind of effort that could have been done anonymously Sure, its something each group training an…

So his contribution consisted of putting Bibliotik into plaintext format? Or is there more to it?

Re: The Battle over Books3

#46
Interesting. When it's copyrighted works in a digital form (plain text, ePub, whatever) it's a legal issue.

ROT13 the text is the data still a copyright violation? I imagine so since it's a trivial thing to restore it to a legally volatile form.

I understand that an unresolved issue is whether, once ingested into an LLM, the trained LLM is in violation of copyright. One wonders if human readers too are in violation of copyright for having been "trained" as well when they read a book.

Is a "brain transplant" from one LLM to another a thing? Perhaps just a trivial copy of the node weights or whatever they're called. That would would not let the target LLM off the hook with regard to copyright violation I expect.

But what if one LLM "taught" another. Maybe that is not a thing yet.

Re: The Battle over Books3

#47
post #38
post #4

Earlier quoted context omitted.

Laws. And the lawsuits that arise when such theft is discovered.

Who is going to enforce those laws? Big publishers are more than happy to settle with AI companies - they just their slice of the pie after all. But who is going to protect, say, your Hacker News comments? Are you going to sue the AI company? Is YC going to sue? Are you going to sue YC for not banning crawlers in their robots.txt?

Why would I even bother doing that? I just write comments, it's not some ineffable wisdom. I'm writing them in public. I don't expect to profit from them somehow.

Heck, collecting whatever cents I might be owed for being a drop in the ocean is a losing move in my country.

Re: The Battle over Books3

#48

From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?

39.52Gb gzipped and 108.79GB uncompressed.

Re: The Battle over Books3

#49
post #34

What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators. Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legisla…

[flagged]

Re: The Battle over Books3

#50
post #49
post #34

What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators. Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legisla…

[flagged]

How so ? Copyright was introduced along with the print press to protect the works of authors. When you buy a copy of a book it's not exactly private property of the author that you buy, is it ?
Post reply on HN