Earlier quoted context omitted.
The question is, "Do the creators of bibliotik have proper licenses to aggregate those books into a dataset?" The dataset is certainly a derived work. Edit: https://the-eye.eu/public/Books/humble_books_20180509/ Cue the "they're depriving me of my income" complaints from authors.
>The dataset is certainly a derived work That was not the question.
Books in .txt format for AI training purposes
31–40 of 97 posts
Re: Books in .txt format for AI training purposes
#32This is very cool and useful, thanks. However, saying: "Suppose you wanted to train a world-class GPT model, just like OpenAI. How? You have no data. Now you do. Now everyone does" Is quite a stretch. You need much, much more that 36GB of quality data to train a world-class GPT model.
Obviously our wetware is optimized for it and we aren't true blank slates, but the enormous magnitude of the discrepancy makes me think we're on the wrong path.
Re: Books in .txt format for AI training purposes
#33Earlier quoted context omitted.
>The dataset is certainly a derived work That was not the question.
" Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned. "
Re: Books in .txt format for AI training purposes
#34Earlier quoted context omitted.
> Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned. Are you objecting to the use of data for training, or to the compilation of training datasets?
> Are you objecting to the use of data for training, or to the compilation of training datasets? The GP is objecting to the casualness with which privacy rights and intellectual property rights are ignored by so many in the AI community, not to a choice of whether they object to one or another specific manifestation of how the community is doing so. [edit added]: Those who would argue "no, there is no casual ignoring…
Re: Books in .txt format for AI training purposes
#35This is very cool and useful, thanks. However, saying: "Suppose you wanted to train a world-class GPT model, just like OpenAI. How? You have no data. Now you do. Now everyone does" Is quite a stretch. You need much, much more that 36GB of quality data to train a world-class GPT model.
The GPT approach must be very poor; with only a fraction of that information you can train a human to perform better than any GPT model shown thus far. When you look at power burned doing it, the gap is even wider. Obviously our wetware is optimized for it and we aren't true blank slates, but the enormous magnitude of the discrepancy makes me think we're on the wrong path.
Instead, as you pointed out, it's just proof of how efficient our wetware is.
Re: Books in .txt format for AI training purposes
#36Re: Books in .txt format for AI training purposes
#37Earlier quoted context omitted.
The question is..if a AI reads a book is it against copyright? Or is the trained model a derived work of those books?
Of course it is a derived work.
Re: Books in .txt format for AI training purposes
#38Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned.
Aside: it's interesting to think about the legal ramifications of copyrights in the face of AI. If my AI model "reads" a bunch of text, is the model now a "derived work" of the copyrighted work? Perhaps. But of course humans don't infringe when they watch/listen/read a copyrighted work and it changes their state of mind. Perhaps the model can be said not to infringe if it sufficiently abstracts concepts of the art in…
It'll work well until we can upload ourselves to the cloud and we'll have to revisit it
Re: Books in .txt format for AI training purposes
#39Always impressive how the AI community, both inside large companies and not, seems to assume data just exists for them to use, copyright or personal rights be damned.
Hell, unless you paid for it, you shouldn’t even have the collection to train on, whether you distribute it or not.
Effectively, everyone’s focusing on whether your story(/model) is sufficiently transformative from Die Hard to be fair use for you to distribute it, without addressing the complaint that you snuck into the theater too, and held the door for others to follow me in.