Live data from Hacker News

The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

dallasexpress.com

61–70 of 79 posts

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#61
post #34
post #31

Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a s…

10 years may be too little but generally I agree. What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle…

far too often people care about things for purpose, not because they care. There are things of equal sorrow that don't get cared for because it serves no purpose. If your cause isn't ammunition for something that people already want to do then it will struggle to gain traction.

For example, nobody cared about data centre water or electricity use when it was serving us up cat pictures. Collectively the past 20 years of doing that is still significantly more use than AI data centres so far. We suddenly care about it because its anti-AI ammunition, not because we care.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#62

Earlier quoted context omitted.

So what are they doing about the online data they ingested?

Copying it. Because they clearly believe that they’re allowed to do so. For physical books, they believe they’re not allowed to, hence the destruction. It’s abhorrent, I agree… maybe if your world domination plan involves destroying rare books, you should change the plan rather than saying “well our hands our tied, the law says we have to”

I mean they used to have a different plan, which involved not destroying the books, and then they got sued for not destroying the books, by us, so now they are destroying the books. What did you think would happen?

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#63
post #60
post #29

Earlier quoted context omitted.

Destroying your copy of a copyrighted work does not violate copyright.

Format shifting is allowed. Making copies isn't. By destroying the original, it's format shifting.

Right. But destroying a book even without making a copy is also allowed. It's very very common, in fact; something approaching a billion books are destroyed each year in the US.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#64
post #24

Earlier quoted context omitted.

The "putting it online" part is where you are f-ing up.

But it's transformed now! Edit(seriously): >Transformativeness is a characteristic of such derivative works that makes them transcend, or place in a new light, the underlying works on which they are based. In computer- and Internet-related works, the transformative characteristic of the later work is often that it provides the public with a benefit not previously available to it, which would otherwise remain unavaila…

It has to be sufficiently transformed. Simply scanning it doesn't qualify. The legal system by contrast is leaning toward training of AI as being sufficiently transformative.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#65

Forever preserving the content of a rare book by consuming one (1) copy is the kind of thing we should always be doing more of. Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.

One (1) company destroys and (illegally?) copies a rare book. This knowledge is now of that company, not of humanity. Other companies will feel the need to do the same and before you know it knowledge is walled off in the gardens of the AI companies, and the rare books are gone and destroyed.

When are you or anyone more likely to get access to a rare book:

a) When a company has made a digitized copy of it

b) When a few previous hard copies are kept somewhere in the world

Idk what Anthropic plans to do with this, but note that my original comment was not about Anthropic at all. I simply think it's a good tradeoff to destroy a few copies of rare books if that helps digitizing them. I am not against simply legally forcing companies that do this to release those copies in due time. In the meantime I am also happy to have them as part of the LLM corpus.

> illegally?

No, copying rare books is very likely not illegal, because, presumably, they are quite old.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#66
post #27

Earlier quoted context omitted.

The scan doesn't go away, and can be released when the copyright expires. Oh no, a lump of cellulose is gone. Will no one think of the fibers?

> can be released when the copyright expires But it ... won't be? The companies have no motivation to do so. The copyright on lots of these books are surely already expired, so if they wanted to they could be putting these up now. I'm sure internally this is viewed as a corpus of knowledge they have that their competitors do not, so they will not release them unless something forces them to.

Yes, without any justification let's assume the worst. This is true good faith arguing.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#67
post #48
post #45

Earlier quoted context omitted.

It's possible to be opposed to excessive copyright and also be opposed to the mass theft of media for profit and subsequent abuse of copyright that the scrapers are guilty of. The AI pillagers are taking everything in, then gatekeeping access. And as they acknowledge with their "nothing after 2022" preferences, they (or their users at least) are salting the earth for future seekers of knowledge.

> It's possible to be opposed to excessive copyright Define excessive. Who makes that determination and how?

There's a substantial economic literature on this topic.

Among other factors: time value of money and future-value discounting mean that economically ever-expanding copyright terms offer virtually no present-value benefit.

Most works see virtually all of their economic value in the first few years of publication. There are some (rare) long-tailed exceptions. The original US copyright term of 14 years, with extensions, actually fits the economics pretty well.

What extreme copyright duration does do is:

1. Create vast copyright-holding monopolies, often of works for which authors were poorly compensated if at all. (For the latter: academic publishing.)

2. Create a vast category of "orphan works" which aren't or cannot be published, in the latter case because rights simply cannot be clearly established, or competing claims (survivors, estates, publishers) exist.

"Seventeen Famous Economists Weigh In On Copyright : the Role of Theory , Empirics , and Network Effects" https://jolt.law.harvard.edu/articles/pdf/v18/18HarvJLTech43...> (2005)

"The true impact of shorter and longer copyright durations: from authors’ earnings to cultural creativity and diversity" Jimmyn Parc & Patrick Messerlin (2020) https://www.tandfonline.com/doi/full/10.1080/10286632.2020.1...>

Orphan Works: https://en.wikipedia.org/wiki/Orphan_work>

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#68
post #62

Earlier quoted context omitted.

Copying it. Because they clearly believe that they’re allowed to do so. For physical books, they believe they’re not allowed to, hence the destruction. It’s abhorrent, I agree… maybe if your world domination plan involves destroying rare books, you should change the plan rather than saying “well our hands our tied, the law says we have to”

I mean they used to have a different plan, which involved not destroying the books, and then they got sued for not destroying the books, by us, so now they are destroying the books. What did you think would happen?

They got sued for violating copyright in general, not for “not destroying books” in particular. Destroying the books was a legal trick they used to continue scanning them without breaking the law.

If I were running the AI company and my lawyers came back with “teeeechnically we can still scan them if we burn them after”, I would reply with “I guess we’re not scanning them”, but I lack the sociopathic instincts required to be a tech CEO I guess.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#69
post #27

Earlier quoted context omitted.

The scan doesn't go away, and can be released when the copyright expires. Oh no, a lump of cellulose is gone. Will no one think of the fibers?

> can be released when the copyright expires But it ... won't be? The companies have no motivation to do so. The copyright on lots of these books are surely already expired, so if they wanted to they could be putting these up now. I'm sure internally this is viewed as a corpus of knowledge they have that their competitors do not, so they will not release them unless something forces them to.

We have a partial list from https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...:

> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

Even if the authors died immediately after publication there is no way the copyright has expired.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#70

There is good discussion on the previous thread from 6 days ago: https://news.ycombinator.com/item?id=49068738

It’s been stunning how often this story is reposted from various news sources across Reddit and Hacker News.

Each time it’s some different outlet - but when you dig in, the piece is just verbal framing around the original story written by 404 media:

https://archive.is/9MQrK

What’s even more stunning is that the original article doesn’t provide evidence that “rare” books are being destroyed. That doesn’t even appear in the original article title.

Post reply on HN