Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

351–360 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#351
The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.

Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."

What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.

Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.

If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?

But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.

Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.

To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?

But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.

I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!

The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.

Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:

a) EU companies making ML models have to self-sabotage against their competition.

b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.

Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.

[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...

[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client

Re: AI companies destroy physical books – let's scan rare books before it's too late

#352

Earlier quoted context omitted.

From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.

They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying > The print original was destroyed. One replaced the other. So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal…

Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#353

You ask "Why destroy physical books?" I ask "Why save physical books?" If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribut…

[flagged]

Re: AI companies destroy physical books – let's scan rare books before it's too late

#354

It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.

The US Library of Congress is already a book depository, i.e., it has a copy of every book published in the United States. Same for the British Library for the UK and Ireland. Similar depositories exist for most other countries who care for their culture.

Which is really why the outrage cycle over Anthropic's actions is largely misplaced.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#355
post #34

The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022." Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire…

Also there was a highly discussed paper talking about how “touched by machines” content will kill llms. About a month after the papers first llms trained with “touched by machines” content appeared an the capabilities of the models got huge upgrade by using that dirty content

Re: AI companies destroy physical books – let's scan rare books before it's too late

#356

You ask "Why destroy physical books?" I ask "Why save physical books?" If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribut…

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!

Re: AI companies destroy physical books – let's scan rare books before it's too late

#357

What often gets missed is that they are buy one physical copy and turning it into a digital copy. They have done zero to destroy the durability. In fact, it’s probably more durable. If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.

> turning it into a digital copy. Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?

The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.

As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#358

You ask "Why destroy physical books?" I ask "Why save physical books?" If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribut…

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.

Same can be said about many books.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#359
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.

Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.

It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.

Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court

Re: AI companies destroy physical books – let's scan rare books before it's too late

#360

It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units. Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity. By the way, I always search for second hand books. Most of the books there ar…

Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
Post reply on HN