We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals). We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in…
Most likely they don't know you exist. If you contacted them first, maybe something could be arranged...
AI companies are shredding rare books
91–100 of 559 posts
Re: AI companies are shredding rare books
#92At the very least why not upload the scanned books to the internet archive while already at it? This is highly disturbing news; is this standard practice? What did Google Books do before?
Re: AI companies are shredding rare books
#93Re: AI companies are shredding rare books
#94It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books: > You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge s…
There is zero reason to shred 18th century books. Any such books are out of copyright.
Re: AI companies are shredding rare books
#95> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright…
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper. What happens to the pages after? No one needs them anymore, so they get mulched and recycled. That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to sc…
In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
Re: AI companies are shredding rare books
#96Re: AI companies are shredding rare books
#97Earlier quoted context omitted.
> why dilute it with this kind of bullshit: > > It is equivalent to book burning in the past. A form of thought control That's only bullshit if you trust AI companies to serve the book contents without alteration.
I don’t trust or expect AI companies to serve books at all. That’s not what they are scanning them for.
The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.
The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Re: AI companies are shredding rare books
#98This is the opposite of a book burning. These books which only a few would ever know the names of, let alone find, let alone read, are being digitised so they can be found in electronic searches.
> being digitised so they can be found in electronic searches You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.
Re: AI companies are shredding rare books
#99The segment that talks about rare books:
> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]
> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
Re: AI companies are shredding rare books
#100> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright…
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper. What happens to the pages after? No one needs them anymore, so they get mulched and recycled. That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to sc…
Because there’d be much less content created in any media to capture in the first place.