Live data from Hacker News

Build a better Bookshelf

huytd.github.io

11–20 of 66 posts

Re: Build a better Bookshelf

#12
post #7

I’ve found books.google.com be alright for this as well.

The author specifically noted: And the search result is limited to just my books, not the whole internet, that's the point of scanning instead of Googling.

And the above commenter specifically noted

I’ve found books.google.com be alright for this as well.

Google Books doesn't index the internet. It indexes books. Including using OCR.

Re: Build a better Bookshelf

#13

The fact that we have to scan any technical book published after 1980 is a distortion of capitalism. The obvious most efficient solution for everyone would be to have the darn fully searchable digital version of the book + the source code.

It's plausible that the fact we have so many amazing technical books is also a "distortion of capitalism". The obvious most efficient course of action is to not write a book.

Re: Build a better Bookshelf

#14
post #7

Earlier quoted context omitted.

The author specifically noted: And the search result is limited to just my books, not the whole internet, that's the point of scanning instead of Googling.

And the above commenter specifically noted I’ve found books.google.com be alright for this as well. Google Books doesn't index the internet. It indexes books. Including using OCR.

books.google.com is mostly every book ever printed. Keeping it to your local bookshelf means you'll actually have access to these books.

Re: Build a better Bookshelf

#16
Does anyone have any experience with book scanning in general? I've been eyeballing a unit from "CZUR" but am a bit skeptical of the product in general. I would prefer to buy something more generic/high end HW wise like a V-shaped scanner where you bring your own DSLRs but can't find out if there is a serious open source software platform for them.

Re: Build a better Bookshelf

#17
post #7

Earlier quoted context omitted.

The author specifically noted: And the search result is limited to just my books, not the whole internet, that's the point of scanning instead of Googling.

And the above commenter specifically noted I’ve found books.google.com be alright for this as well. Google Books doesn't index the internet. It indexes books. Including using OCR.

Consider for a moment the critical difference between my books on my bookshelves and Google-indexed books in the ether.

Re: Build a better Bookshelf

#19
post #16

Does anyone have any experience with book scanning in general? I've been eyeballing a unit from "CZUR" but am a bit skeptical of the product in general. I would prefer to buy something more generic/high end HW wise like a V-shaped scanner where you bring your own DSLRs but can't find out if there is a serious open source software platform for them.

https://www.diybookscanner.org/

Re: Build a better Bookshelf

#20
post #16

Does anyone have any experience with book scanning in general? I've been eyeballing a unit from "CZUR" but am a bit skeptical of the product in general. I would prefer to buy something more generic/high end HW wise like a V-shaped scanner where you bring your own DSLRs but can't find out if there is a serious open source software platform for them.

I made a hardware DIY Book scanner and run it off a Rpi using https://github.com/Tenrec-Builders/pi-scan with two Nikon J1 cameras. If you're not an experienced handy-person be prepared for frustration and spending more than you think on tools.

Otherwise very happy with the result and experience. I can scan about 800 pages per hour currently. Once scanned I use scan tailor to adjust the pages then a small script to tesseract OCR them before creating a PDF.

I typically do all the scanning while watching Netflix as it doesn't require a lot of attention once you get in the flow.

Post reply on HN