Live data from Hacker News

A 'bananas' order for 5000 obscure book titles fuels suspicion

irishtimes.com

51–60 of 76 posts

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#51
post #36

> with a focus only on books with an ISBN number, the identification system introduced in 1966 So it's not likely to be rare and precious books. It's things which are already in the Library of Congress or its many equivalents. If they're forced by laws to destroy the results of the scans after training on them, as some have implied, that's bad, on the chance there is some actual lost media in there. But if they keep…

SBN started in 1967 in UK (W.H.Smith), official international book numbering (ISBN) is after 1970. Source: https://www.isbn.org/ISBN_history

Strange to think of WH Smith as being at the forefront of technology.

Also weird that in the space of 5 years, they'd gone from an idea, got national buy in, and then got an international standard sorted.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#52
post #25

Earlier quoted context omitted.

Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

> Anyone that wants to train a model needs to procure and destroy their own physical copy of each book! Why? Couldn't they resell or give away the books after scanning them?

chop off the spine and you can use an automated scanner. to keep the book intact you'll need a flatbed scanner and an operator.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#53
post #25

Earlier quoted context omitted.

Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

> Anyone that wants to train a model needs to procure and destroy their own physical copy of each book! Why? Couldn't they resell or give away the books after scanning them?

Per the courts, it actually helps their "fair use" case if they destroy the book instead of re-selling it.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#54
post #25

Earlier quoted context omitted.

Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

How is this in any way, shape, or form a promotion of the arts and sciences anymore? This could very easily be turned into a preservation and archiving operation with just a tweak of the laws, or a carve-out. And make the bank once, and make it legal to train on? How people in the bank get compensated is a different question, but -while almost impossible to settle on an individual basis- could be settled reasonably i…

How is this in any way, shape, or form a promotion of the arts and sciences anymore?

Well, if the book is still in production the author will get the royalties. Though if it were still an active release they should just sell digital copies.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#55
Two American self published nonfiction authors I know reported large purchases of new books earlier this year, more than 50 units at a time, via Amazon. This is unusual because the books are not well known. Because the books do not sell well otherwise, they were very few used copies available for sale on Amazon.

Here's my AI theory: someone is purchasing lots of books en masse to be packaged and sold for LLM training to multiple clients (not just a single company like Anthropic), but only scanning once and then reselling the scans while preserving a "chain of custody" proof that individual copies were purchased. Claude shot that idea down for reasons related to the intricacies of US first sale doctrine and copyright law, but had an ambiguous response when I proposed that it might be Chinese companies doing something similar for training AI models that are intended for possible resale to overseas markets.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#56
post #9

Earlier quoted context omitted.

As long as the AI devouring a book isn't making it harder for a person to get access to a copy of the book. There might not be many copies of some out-of-print books.

Would be nice to at least force them to open their scanned books data, after all they stole everything else, they can contribute some back

And what, potentially make things easier for a competitor? Not in capitalism sir

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#58
post #30
post #6

Is this a second form of AI psychosis, as in AI training corpus psychosis? The need for all the content, Moar!! Feed me. Is this the Paperclip Maximiser in the form of the Training Material Maximiser? Will it be that, in the end, the lack of "Pass Your Driving Test, 2018 Edition" was the cause of driving rule hallucinations in all previous models? Will this finally get OpenAI back in front of Anthropic? Hurrah! We fo…

> The need for all the content, Moar!! Feed me "MOAR input!!!" - Johnny 5, Short Circuit. That was a great film. Seems it was unexpectedly prescient too.

https://www.youtube.com/watch?v=WnTKllDbu5o

As a teenager I remember I was amazed at the speed at which Johnny 5 would consume that book.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#59

Earlier quoted context omitted.

Would be nice to at least force them to open their scanned books data, after all they stole everything else, they can contribute some back

"Open" it how, exactly? Make it available for free? As in violate copyright?

Perhaps Anna's Archive will be the last bastion of some books, especially if they manage to acquire the internal collections of some of the AI companies, à la Spotify. (ie. the parts of the collections that didn't come from Anna in the first place.)
Post reply on HN