Does anyone have more content about what makes Books3 so special relative to Bibliotik? Was it processed somehow, or just compiled into a single file? I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that…
They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.)
> This is the kind of effort that could have been done anonymously
Sure, its something each group training an AI could do independently at greater aggregate cost until someone succeeds in taking the original source down, but not only would that be costlier, but it in would involve less transparency and comparability across model architecture, or at least required the transparent, comparable trained version to be different from the full version.