Live data from Hacker News

Books in .txt format for AI training purposes

twitter.com

31–40 of 97 posts

Re: Books in .txt format for AI training purposes

#31
post #30
post #29

Earlier quoted context omitted.

The question is, "Do the creators of bibliotik have proper licenses to aggregate those books into a dataset?" The dataset is certainly a derived work. Edit: https://the-eye.eu/public/Books/humble_books_20180509/ Cue the "they're depriving me of my income" complaints from authors.

>The dataset is certainly a derived work That was not the question.

"Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned."

Re: Books in .txt format for AI training purposes

#32

This is very cool and useful, thanks. However, saying: "Suppose you wanted to train a world-class GPT model, just like OpenAI. How? You have no data. Now you do. Now everyone does" Is quite a stretch. You need much, much more that 36GB of quality data to train a world-class GPT model.

The GPT approach must be very poor; with only a fraction of that information you can train a human to perform better than any GPT model shown thus far. When you look at power burned doing it, the gap is even wider.

Obviously our wetware is optimized for it and we aren't true blank slates, but the enormous magnitude of the discrepancy makes me think we're on the wrong path.

Re: Books in .txt format for AI training purposes

#33
post #31
post #30

Earlier quoted context omitted.

>The dataset is certainly a derived work That was not the question.

" Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned. "

And you where thinking, lets rephrase that question because it's not already boring and obvious enough?

Re: Books in .txt format for AI training purposes

#34
post #5
post #3

Earlier quoted context omitted.

> Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned. Are you objecting to the use of data for training, or to the compilation of training datasets?

> Are you objecting to the use of data for training, or to the compilation of training datasets? The GP is objecting to the casualness with which privacy rights and intellectual property rights are ignored by so many in the AI community, not to a choice of whether they object to one or another specific manifestation of how the community is doing so. [edit added]: Those who would argue "no, there is no casual ignoring…

Sweet jesus those instructions go on forever, imagine performing that (without breaking into laughter)

Re: Books in .txt format for AI training purposes

#35

This is very cool and useful, thanks. However, saying: "Suppose you wanted to train a world-class GPT model, just like OpenAI. How? You have no data. Now you do. Now everyone does" Is quite a stretch. You need much, much more that 36GB of quality data to train a world-class GPT model.

The GPT approach must be very poor; with only a fraction of that information you can train a human to perform better than any GPT model shown thus far. When you look at power burned doing it, the gap is even wider. Obviously our wetware is optimized for it and we aren't true blank slates, but the enormous magnitude of the discrepancy makes me think we're on the wrong path.

I'm sure it's much more probable than we're in the wrong path, rather than the good one, but I don't think the amount of information needed by a GPT model can give insights to that.

Instead, as you pointed out, it's just proof of how efficient our wetware is.

Re: Books in .txt format for AI training purposes

#37
post #8
post #6

Earlier quoted context omitted.

The question is..if a AI reads a book is it against copyright? Or is the trained model a derived work of those books?

Of course it is a derived work.

I read a Java book and created programs using knowledge given in the book. Is my program a derived work?

Re: Books in .txt format for AI training purposes

#38
post #2

Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned.

Aside: it's interesting to think about the legal ramifications of copyrights in the face of AI. If my AI model "reads" a bunch of text, is the model now a "derived work" of the copyrighted work? Perhaps. But of course humans don't infringe when they watch/listen/read a copyrighted work and it changes their state of mind. Perhaps the model can be said not to infringe if it sufficiently abstracts concepts of the art in…

I'd imagine one day this will be determined by whether you can copy the resulting 'brain' to another system

It'll work well until we can upload ourselves to the cloud and we'll have to revisit it

Re: Books in .txt format for AI training purposes

#39
post #2

Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned.

Seriously. At one point I curated a similar dataset to OP, but because it was largely worked to which I didn’t own the copyright, there’s no way in hell I was going to distribute it to others (because that would be doing what that internet library got sued for). I can’t freely distribute the last 20 years of TV and movies just because it’s useful for academic research.

Hell, unless you paid for it, you shouldn’t even have the collection to train on, whether you distribute it or not.

Effectively, everyone’s focusing on whether your story(/model) is sufficiently transformative from Die Hard to be fair use for you to distribute it, without addressing the complaint that you snuck into the theater too, and held the door for others to follow me in.

Post reply on HN