Live data from Hacker News

A 'bananas' order for 5000 obscure book titles fuels suspicion

irishtimes.com

41–50 of 76 posts

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#41

> with a focus only on books with an ISBN number, the identification system introduced in 1966 So it's not likely to be rare and precious books. It's things which are already in the Library of Congress or its many equivalents. If they're forced by laws to destroy the results of the scans after training on them, as some have implied, that's bad, on the chance there is some actual lost media in there. But if they keep…

There are plenty of rare and precious books with ISBNs.

Rare, certainly: there are plenty of vanity publications that were chucked into the bin by almost everyone who was unlucky enough to be given a copy. Precious? Well, with the help of an electronic friend I found ISBN 978-3-8365-7349-8. Apparently a copy of that book is worth about £40k. Can anyone beat that?

(EDIT: I'm assuming "precious" means the same as valuable here to make the question easier to answer. In fact, of course, people usually say "precious" when they mean a personal or emotional attachment or cultural importance rather than economic value.)

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#42
post #25
post #12

Earlier quoted context omitted.

> But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest. only if they share the scans/transcripts and don’t hide it away.

Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

> Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

Why? Couldn't they resell or give away the books after scanning them?

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#43
post #25
post #12

Earlier quoted context omitted.

> But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest. only if they share the scans/transcripts and don’t hide it away.

Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

How is this in any way, shape, or form a promotion of the arts and sciences anymore?

This could very easily be turned into a preservation and archiving operation with just a tweak of the laws, or a carve-out.

And make the bank once, and make it legal to train on? How people in the bank get compensated is a different question, but -while almost impossible to settle on an individual basis- could be settled reasonably in bulk by some form of mandatory implied statutory contract?

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#44
post #25

Earlier quoted context omitted.

Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying). Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

> Anyone that wants to train a model needs to procure and destroy their own physical copy of each book! Why? Couldn't they resell or give away the books after scanning them?

It's less work if you use a destructive method. They rip the spine off to get individual flatter sheets instead of scanning a bound book.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#45
I wish we had competent regulators. Force the scans to be purchased through an official repository. Save the scans and sell them to other companies who also want to train models. Force the companies to make all book requests through the official repository. Give some portion of money paid back to the publishers. This is a solvable problem.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#46

When I hear these stories of AI companies buying all the books, I think back to Kevin Kelly in 2008 or 2009. He talked about how books are less expensive than at any other time in history and easier to buy and that could change so it makes sense to buy a lot of books. > [...] I was near to the point of actually digitizing and getting rid of all my paper books. > I was that close about five years ago, but then I had a…

[flagged]

The second paragraph in that story lays out the main bullets.

He's worth paying attention to. For the past thirty years or so that I've been aware of him he's been more prescient than most.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#47

> with a focus only on books with an ISBN number, the identification system introduced in 1966 So it's not likely to be rare and precious books. It's things which are already in the Library of Congress or its many equivalents. If they're forced by laws to destroy the results of the scans after training on them, as some have implied, that's bad, on the chance there is some actual lost media in there. But if they keep…

There are plenty of rare and precious books with ISBNs.

In the lost media sense? I'm not sure.

Of course there are always post 1960 books which are rare and valuable for being very limited run first editions, signed copies etc.

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#49
post #48

Can't these genius AI companies spend a few pennies to develop a non-destructive book scanner? What's the use of all this AI if they can't figure out this relatively simple task.

no, because that's cost and everything needs to be the cheapest possible. only valid spend is money going to the shareholders. and chopping off the spine and putting the stack of paper into an automated scanner saves sooo much money over a $5/hr employee flipping pages all day

i think that they might actually want to do the ram gambit all over again - buy up these rare books so that competitors can't get them. destroying them after scanning makes doubly sure that the competition won't get their hands on it

Re: A 'bananas' order for 5000 obscure book titles fuels suspicion

#50
post #48

Can't these genius AI companies spend a few pennies to develop a non-destructive book scanner? What's the use of all this AI if they can't figure out this relatively simple task.

My understanding was that they had to legally destroy the books after they processed them for whatever copyright strangeness was involved.
Post reply on HN