Earlier quoted context omitted.
I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?
They aren’t trained on compressed plaintext so I’m not sure of the relevance there. But regardless it’s my understanding that’s modern models are trained with orders of magnitude more storage than their parameters require. But it’s possible I’m incorrect. This is getting to the fringe of my knowledge of concrete LLM details.
AI companies are shredding rare books
451–460 of 559 posts
Re: AI companies are shredding rare books
#452Earlier quoted context omitted.
If it's copyright is expired why not?
The books in question are not antique. The Twitter post makes this claim but so far as I can tell it’s not based in fact. 404 Media published a story about this as well and cites a bookseller who notes that all of the books there have sold have had ISBNs (and are thus from 1967 or later and generally would have active copyright). ”very large purchases were of books that had little in common, except for the fact that…
Re: AI companies are shredding rare books
#453Earlier quoted context omitted.
I have even less sympathy for IP stealing LLM operators
It’s been determined that training on lawfully acquired works is fair use. Presumably in this discussion of shredding physical books Dario and Sam are not pulling heists at the local library. I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/... Similar conclusion vs met…
Re: AI companies are shredding rare books
#454Earlier quoted context omitted.
20 years to make some money, and then we set the work free for the public benefit. If it's good enough for patents, I don't see why it isn't good enough for copyrights.
For corporations sure. For individual authors that's certainly not fair. Especially since it makes it easier for corporations to exploit their work without paying them anything. > If it's good enough for patents, I don't see why it isn't good enough for copyrights Because there are fundamentally differ concepts and serve different purposes?
The purposes of both are: "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
Like, this is a made up regime with a specific intent. The fact that we treat copyrights and patents differently is an accident of history. I think we could quite reasonably choose a different period of time (and in fact, have done so several times over the past few hundred years) and still promote progress.
I think it's very reasonable to say that one good idea should not be enough to let you coast your whole life, you should be prodded to cough up 3 good ideas. Further, it reduces corporate power at the other end by allowing individuals to play in coroporate properties after a relatively short time. You could be futzing around with, idk, a copyright free Cars under my proposed regime.
Re: AI companies are shredding rare books
#455Re: AI companies are shredding rare books
#456Earlier quoted context omitted.
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper. What happens to the pages after? No one needs them anymore, so they get mulched and recycled. That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to sc…
Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.
Re: AI companies are shredding rare books
#457Earlier quoted context omitted.
Just to be clear, publishers hadn't accepted the "controlled digital lending" (CDL) premise, not even with the one-to-one ratio. Their position was always "first sale ends when the atoms do". There was even controlling precedent: a few years before IA tried their online lending library thing, there was an "MP3 resale" company called ReDigi that had lost on very similar grounds. The publishers suing IA even made sure…
The big insanity is tying this to AI. Shredding books is about format shifting; it's a concession hard-won from copyright establishment, which would otherwise be more than happy to deny you the option to convert the media you owned from physical to digital. AI training happens to be one of the fields exercising that option, but since it's the current favorite topic for people to hate on, here we are.
The reason why AI companies don't do this is that they're cheap and desperate for training tokens. Same reason why they have scrapers that will happily overload web interfaces for Git repos following links to everything, even though you can just Git clone the repo with far less stress on the host. The AI people are ultimately there just to pillage as much knowledge as they can as fast as possible. Their scraping practices are slap-dash garbage.
[0] https://ones-and-zeroes.ghost.io/scanning-all-the-books-the-...
Re: AI companies are shredding rare books
#458Earlier quoted context omitted.
Then just make it the rule that the copyright expires after 30 years or when the author dies, whichever comes last.
Why not just 30 years? Patents get a flat 20. Not that I'm arguing for 30 per se, just that I don't see what goals of copyright would be advanced more by adding an "or until death" complication.
That's not the same for a work of fiction or a piece of music. Case in point apparently books sales for the Odyssey are massively up - when it was originally written in 7-8 BC :-)
Also most books etc don't make much, if any money - an publisher/author might rely on a the cummulative effect of a number of revenue streams built over time.
Also the effect of exclusivity is different - for patents you are potentially blocking the area of innovation you have patented by your exclusivity.
That's not the same societal effect as somebody not being able to copy mickey mouse.
So they aren't exactly the same - however I'm not proposing a 3000 year copyright :-)
Re: AI companies are shredding rare books
#459I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…
Copyright can be intrinsic, sure, but only to maybe 5-10yrs, far less than the maximum to incentivize registration (and thus incentivizing archiving of culture).
I should be able to freely download every Disney film prior to 2006, compile Killer7 for the fun of it, buy a hardback of the LOTR trilogy from any printer I wish, and do so without any interference from the copyright holders.
However, because we don’t live in a utopia and Sonny Bozo became a politician, I’ll be waving a certain jolly flag for the foreseeable future.
Re: AI companies are shredding rare books
#460Earlier quoted context omitted.
Either that or some exponentially increasing tax so that Disney can keep their vault. (I'm perfectly fine with them keeping it if they pay some proper taxes.)
They pay taxes every time that they make money off Mickey Mouse, regardless of copyright status.