Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

801–810 of 962 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#801
post #773
post #729

Earlier quoted context omitted.

I think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually nece…

If it isn't scanned destructively, is it a liability? Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain? These questions imply that there's a liabil…

Questions don't imply anything, you're just asking things. I don't have the answers to your questions, but the point is that Google was able to scan books without destroying them while still avoiding legal consequences (with some effort).

Re: AI companies destroy physical books – let's scan rare books before it's too late

#802
post #788

Earlier quoted context omitted.

Wrong end of the pipeline; we should instead demand digital copies of media be sent to the Library of Congress in order to obtain copyright, along with a registration fee to pay for indefinite storage and other costs. Registration should be mandatory if you want copyright. For things like books where a machine readable text format existed, it should be mandatory to include (so no requiring OCR). Access to the archive…

That's already the case, though the "digital rather than physical" as a preference could be something that legislation would improve. https://www.copyright.gov/mandatory/ > All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17). > This law requires two copies of each work published in the United States…

It is not. We currently give ~infinite copyright automatically for nothing in return:

> Neither the deposit requirements of this subsection nor the acquisition provisions of subsection (e) are conditions of copyright protection.

https://www.copyright.gov/title17/92chap4.html#407

Re: AI companies destroy physical books – let's scan rare books before it's too late

#804
post #662

Earlier quoted context omitted.

The Library of Congress already has a copy of every book published in the US. How would this help?

If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.

> it would be easier for them to distribute the work after the copyright of the work expires

Copyright does not expire for a very long time. Harry Potter and the Sorcerer's Stone was released ~30 years ago in 1997. It remains protected for the duration of the life of the author (J.K. Rowling) plus 70 years.

Given actuarial tables from the UK[1], this works out to be around ~95 years from now (~2120).

[1] https://www.ons.gov.uk/peoplepopulationandcommunity/birthsde...

Re: AI companies destroy physical books – let's scan rare books before it's too late

#805

Earlier quoted context omitted.

Then at least there would be outrage to drive the passing of the needed legislation which otherwise hasn't come to pass anyway.

Copyright law desperately needs a production requirement or allowance. The copyright owner must make new copies of the work available; the price must be no greater than the original price (not inflation adjusted). And if they fail to do so, anyone may produce copies and escrow the original price (less the cost of production) for collection by the copyright holder. That means that orphan works are effectively in the p…

As a photographer, do I have to make every photograph that I've ever sold available to anyone to buy forever more? Can I refuse to sell a print to someone? I wasn't famous when I sold one for $20 back in the 90s... if I became famous, would I still need to sell that at $20 (inflation adjusted)?

What happens to limited editions of print runs? Can I not make a run of 200 prints anymore because the 201st will be something that someone could request?

Does a musician have to license any song they made to anyone who asks? Can they refuse to license a song to some organization they disagree with and not have it fall into the orphan works category?

---

Amending copyright to the way you describe requires a renegotiation of the TRIPS agreement ( https://en.wikipedia.org/wiki/TRIPS_Agreement ) with all the nations of the WTO (or withdrawing from the WTO).

Re: AI companies destroy physical books – let's scan rare books before it's too late

#806

Earlier quoted context omitted.

They could use them as training data, without providing access the actual books.

At least we would all benefit from the books this way, so long as legal nonsense keeps the scans unavailable to the public.

I don’t think locking the content of rare books away in the hands of corporations who only give us access to tools trained on the books, and not the actual book, is a good path forward.

This doesn’t incentivize them to be good stewards of this data and making anything in the public domain available. It incentivizes less access to the source material, having to blindly trust their tools, and is effectively automating plagiarism.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#807
post #788

Earlier quoted context omitted.

That's already the case, though the "digital rather than physical" as a preference could be something that legislation would improve. https://www.copyright.gov/mandatory/ > All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17). > This law requires two copies of each work published in the United States…

It is not. We currently give ~infinite copyright automatically for nothing in return: > Neither the deposit requirements of this subsection nor the acquisition provisions of subsection (e) are conditions of copyright protection. https://www.copyright.gov/title17/92chap4.html#407

Copyright protection is automatic.

Publishing of copyrighted material requires that it be deposited with the Library of Congress.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#808
post #662

Earlier quoted context omitted.

If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.

> it would be easier for them to distribute the work after the copyright of the work expires Copyright does not expire for a very long time. Harry Potter and the Sorcerer's Stone was released ~30 years ago in 1997. It remains protected for the duration of the life of the author (J.K. Rowling) plus 70 years. Given actuarial tables from the UK[1], this works out to be around ~95 years from now (~2120). [1] https://www.…

Certainly. But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.

The library of congress does distribute some digitized works that are out of copyright. And it does digitize some works for archival and distribution, but having additional works digitized for (eventual) public use could be nice.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#809
post #807

Earlier quoted context omitted.

It is not. We currently give ~infinite copyright automatically for nothing in return: > Neither the deposit requirements of this subsection nor the acquisition provisions of subsection (e) are conditions of copyright protection. https://www.copyright.gov/title17/92chap4.html#407

Copyright protection is automatic. Publishing of copyrighted material requires that it be deposited with the Library of Congress.

Right, that's why I said we should make copyright require deposit and registration (like it used to). Publish your work without its copyright ID for people to use to reference the LoC database? It is now public domain.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#810

Earlier quoted context omitted.

Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans? How does this help anything, except create more work to throw in the trash?

This is already required for new books published in the US, and has been for more than one hundred years. It’s called “mandatory deposit”

Case law seems to be that mandatory deposit is unconstitutional, fwiw.

https://en.wikipedia.org/wiki/Valancourt_Books_v._Garland

Post reply on HN