Live data from Hacker News

Books in .txt format for AI training purposes

twitter.com

21–30 of 97 posts

Re: Books in .txt format for AI training purposes

#21
post #4

is bibliotik db bigger than libgen? any tips how to get invited?

I am not sure if bibliotik has more books than libgen by sheer numbers but I am pretty sure that bibliotiks books are of higher quality and better organized at the very least. Oh and if you haven't already got your foot in the door or know somebody who does it is unlikely that people like you and I will simply get invited by some kind stranger on the internet.

Also, a lot of books that were behind the bibliotik signup wall have been freed by /u/-Archivist.

https://www.reddit.com/r/opendirectories/comments/f2teym/pro...

I personally recommend you start with MyAnonamouse.org. They hold invite interviews twice a week over IRC and I have heard that the community is really welcoming to new users.

https://www.myanonamouse.net/inviteapp.php

Re: Books in .txt format for AI training purposes

#22
post #8
post #6

Earlier quoted context omitted.

The question is..if a AI reads a book is it against copyright? Or is the trained model a derived work of those books?

Of course it is a derived work.

Maybe not so obvious. If you read a book, are you a derivative work?

Re: Books in .txt format for AI training purposes

#23
What is ".txt format"? I'm genuinely curious what they went with here, but don't want to download 36 GB to find out. There isn't a .txt standard, is there?

I would hope for UTF-8, but given the old-school, Windows-y .txt extension, it could just as easily be Windows-1252 or something. LF or CRLF line endings?

Re: Books in .txt format for AI training purposes

#24

There is always Project Gutenberg [1]. I suppose there may be pros and cons to each, with Project Gutenberg being generally older books. [1]. https://www.gutenberg.org/ Fun trivia: I once took an NLP class and discovered that I could predict whether a text was written by Edgar Allan Poe or H. G. Wells with ~86% accuracy based on the presence of one word: "whereupon". The language had changed enough in 60 years such t…

I’ve noticed that “presently” (in the meaning “soon”) is one such indicator for post-WWII English.

British novels written in the 1950s consistently use “presently” to indicate a short passage of time. In the following decades it seems to have been removed from editors’ style guides and replaced mostly by plain old “soon”.

Re: Books in .txt format for AI training purposes

#25
post #6
post #5

Earlier quoted context omitted.

> Are you objecting to the use of data for training, or to the compilation of training datasets? The GP is objecting to the casualness with which privacy rights and intellectual property rights are ignored by so many in the AI community, not to a choice of whether they object to one or another specific manifestation of how the community is doing so. [edit added]: Those who would argue "no, there is no casual ignoring…

The question is..if a AI reads a book is it against copyright? Or is the trained model a derived work of those books?

It is probably not a derived work. To be a derived work it is not enough to simply use the original work in its construction; it must contain major copyrightable elements of the original work. The weights and parameters present in the model do not seem to fit that description; generally you can’t copyright a bunch of numbers. That said, I don’t think this has been confirmed by a major court case yet, and I wouldn’t be surprised if that happens at some point.

Re: Books in .txt format for AI training purposes

#27

This is very cool and useful, thanks. However, saying: "Suppose you wanted to train a world-class GPT model, just like OpenAI. How? You have no data. Now you do. Now everyone does" Is quite a stretch. You need much, much more that 36GB of quality data to train a world-class GPT model.

No, You don't need more. This is compressed text. All of English Wikipedia is roughly 20 GB. You can train amazing and state of the art text models with this data in addition.

Openai states themselves that they trained on "40GB of internet text"

https://openai.com/blog/better-language-models/

Re: Books in .txt format for AI training purposes

#29
post #6
post #5

Earlier quoted context omitted.

> Are you objecting to the use of data for training, or to the compilation of training datasets? The GP is objecting to the casualness with which privacy rights and intellectual property rights are ignored by so many in the AI community, not to a choice of whether they object to one or another specific manifestation of how the community is doing so. [edit added]: Those who would argue "no, there is no casual ignoring…

The question is..if a AI reads a book is it against copyright? Or is the trained model a derived work of those books?

The question is, "Do the creators of bibliotik have proper licenses to aggregate those books into a dataset?" The dataset is certainly a derived work.

Edit: https://the-eye.eu/public/Books/humble_books_20180509/

Cue the "they're depriving me of my income" complaints from authors.

Re: Books in .txt format for AI training purposes

#30
post #29
post #6

Earlier quoted context omitted.

The question is..if a AI reads a book is it against copyright? Or is the trained model a derived work of those books?

The question is, "Do the creators of bibliotik have proper licenses to aggregate those books into a dataset?" The dataset is certainly a derived work. Edit: https://the-eye.eu/public/Books/humble_books_20180509/ Cue the "they're depriving me of my income" complaints from authors.

>The dataset is certainly a derived work

That was not the question.

Post reply on HN