Live data from Hacker News

DIY Book Scanner

diybookscanner.org

31–40 of 132 posts

Re: DIY Book Scanner

#31
post #23

I have most of the entire collection of hardback national geographics from 1930 to 1970. Wonder how legal it would be to scan them. Always wondered.

> from 1930 Good news: It's very likely that the copyright has expired. If you were to scan them, remember to upload them to archive.org for everyone else to see. Bad news: It's only the case if the copyright hasn't been renewed by the owner. Usually most owners don't renew them, but to determine whether or not this is the case, you need to go through huge catalogs of registered entries from the U.S. copyright office…

Copyright is life of author + 70 years. No renewal necessary. Is that not true?

EDIT: No, guess it's not always true. That's only for post 1978. https://www.copyright.gov/help/faq/faq-duration.html

Re: DIY Book Scanner

#32

I built one of these out of pine 2x4s and plywood. I thought it would be cheaper than buying one (I was wrong) but I'm also not a skilled woodworker and had to buy most of the tools. It works quite well and I digitised dozens of textbooks I'd purchased and needed to reference but couldn't carry around every day while finishing my masters. My one had 2 Nikon mirrorless cameras controlled via Pi-Scan. https://github.co…

The dual camera is a design choice I hadn't thought of. I've thought of scanning a couple books, and that's probably the trick for me. Though maybe I'll rotate the book and scan / rotate images separately.

Re: DIY Book Scanner

#33
post #19

I'm very interested in getting into archival (getting started this month after a few more conversations). Your buy button[0] is broken. You're potentially missing out on a few sales due to this. Is 2x 4GB SD card sufficient for your purposes? I've been quoted 50MB TIFF images as a standard, and a lot of books wouldn't fit without swapping SDs at that size. [0] http://store.diybookscanner.org/

archiving what? just curious.

I want to digitize the entire linguistic and spoken corpus of a critically endangered language[0] and convert it to a searchable format to aid in language revival, academic research, and ensuring that an informed debate can occur when the modern usages of the language differ from traditional usages of the language.

Most of the printed books are scattered, but available, but it's akin to an iceberg: there's a significant amount of 'submerged' knowledge about the language in written manuscripts and recorded audio, and this is where a lot of the value comes from. Printed texts are primarily religious, and getting the colloquial usages of words and phrases is very useful.

Many manuscripts aren't digitized at all, or are available and need transcription.

The language is relatively well-recorded (dating back to at least the late 16th century in written form), and yet small enough that a comprehensive reference is viable: estimates of about 5MM words crop up, but even 3x could easily fit in memory on a Digital Ocean droplet, even if fully POS tagged[1]. Texts are also mostly in the public domain, and there's a lot of bilingual texts (which act as a Rosetta Stone).

[0] https://en.wikipedia.org/wiki/Manx_language#Revival

[1] https://en.wikipedia.org/wiki/Part-of-speech_tagging

EDIT: More than happy to talk in depth about this if anyone wants, via comments, or email on my profile.

Re: DIY Book Scanner

#34
post #23

I have most of the entire collection of hardback national geographics from 1930 to 1970. Wonder how legal it would be to scan them. Always wondered.

> from 1930 Good news: It's very likely that the copyright has expired. If you were to scan them, remember to upload them to archive.org for everyone else to see. Bad news: It's only the case if the copyright hasn't been renewed by the owner. Usually most owners don't renew them, but to determine whether or not this is the case, you need to go through huge catalogs of registered entries from the U.S. copyright office…

Incorrect. It is almost certainly still in copyright in the United States. Anything published after 1925 will be in copyright except for those published without notice 1926-1977 or were not renewed 1926-1963. The exceptions almost certainly don't apply to NatGeo.

Scanning typically falls under fair use, so copyright only applies to distribution of the scans.

edit: Maybe you edited or maybe I'm just dumb. Anyway, the problem with relying on a lack of renewal is that you have to prove a negative. NYPL among others have been doing interesting work on this problem: https://www.nypl.org/blog/2019/09/01/historical-copyright-re...

Re: DIY Book Scanner

#35
post #34

Earlier quoted context omitted.

> from 1930 Good news: It's very likely that the copyright has expired. If you were to scan them, remember to upload them to archive.org for everyone else to see. Bad news: It's only the case if the copyright hasn't been renewed by the owner. Usually most owners don't renew them, but to determine whether or not this is the case, you need to go through huge catalogs of registered entries from the U.S. copyright office…

Incorrect. It is almost certainly still in copyright in the United States. Anything published after 1925 will be in copyright except for those published without notice 1926-1977 or were not renewed 1926-1963. The exceptions almost certainly don't apply to NatGeo. Scanning typically falls under fair use, so copyright only applies to distribution of the scans. edit: Maybe you edited or maybe I'm just dumb. Anyway, the…

[deleted]

Re: DIY Book Scanner

#36

Are any of you part of book scanner clubs that might have a database of word counts of famous fiction books? I've found several lists online but it's not a wide selection of books - I'd imagine book scanners might have more. I'd be happy to share the database I've cobbled together.

I briefly participated in an eBookz scene group at the turn of the century, although we didn't keep track of any word counts (nor did we OCR) and we focused on non-fiction, mainly automotive repair manuals. I doubt it's a statistic that the scanners (people/groups) pay attention to.

Re: DIY Book Scanner

#38
UCL-CS had one of these which was deployed in conjunction with the British Library. This is when high pixel count CCDs were super expensive back in the 1980s. Amazing device.

Re: DIY Book Scanner

#39
post #30

My first job out of college was scanning books for the Internet Archive down in the basement of the Library of Congress. Their scanning machines used a foot pedal to raise and lower the glass Platen, so I'd use one hand to flip the page and wiggle the cradle to get things nice and flat and the other would snap the photo. You can get pretty fast after a while, but boy is it mindless. Older books that had been rebound…

Wow that's awesome! I take it you're responsible for a chunk of the books available now on openlibrary.org?

When scanning books like that did you ever see anything interesting or are you so zoned out you don't really pay attention?

Re: DIY Book Scanner

#40

I built one of these out of pine 2x4s and plywood. I thought it would be cheaper than buying one (I was wrong) but I'm also not a skilled woodworker and had to buy most of the tools. It works quite well and I digitised dozens of textbooks I'd purchased and needed to reference but couldn't carry around every day while finishing my masters. My one had 2 Nikon mirrorless cameras controlled via Pi-Scan. https://github.co…

The dual camera is a design choice I hadn't thought of. I've thought of scanning a couple books, and that's probably the trick for me. Though maybe I'll rotate the book and scan / rotate images separately.

It let me capture the pages with the correct orientation and the cameras have a fixed focus on the Platen so it works really well. Then Scantailer can crop automagically and deal with the rest.
Post reply on HN