Live data from Hacker News

Pdfsandwich

tobias-elze.de

21–30 of 65 posts

Re: Pdfsandwich

#22
Years ago I tried tesseract and similar tools to do this but switched to using ABBYY FineReader because the OCR accuracy was way higher and into "usable". Can anyone offer a recent comparison?

Re: Pdfsandwich

#26
I've been a fan of k2pdfopt [1] for years. Single binary, command line + optional gui. Lets you slice and optimize pdfs in any imaginable way, e.g. for e-readers' screens. I think it also does what pdfsandwich does, if I understand things correctly [2].

That said, pdfsandwich's 'one thing well' approach does have an appeal. I will definitely try it out, thanks for posting. Something in its "logo" reminded me of the OpenBSD fish. :)

1: https://www.willus.com/k2pdfopt/

2: https://www.willus.com/k2pdfopt/help/ocr.shtml

Re: Pdfsandwich

#27
post #7

Not to get all technical, but the "Cube Rule of Food" would classify this project as "pdftoast", not "pdfsandwich": https://cuberule.com/

I'd classify it more as a salad. https://saladtheory.github.io/ (some very interesting content there, and includes a section refuting the cube rule)

Re: Pdfsandwich

#29
i have used this extensively and i love this software. you just supply a file and it crunches the numbers and you get an output.

there is something called scantailor if you are scanning books yourself. that gives you more contrl over the orientation and margins and ocr and contrast controls. that said, pdfsandwich gives you a miniscule file size which i could not achieve otherwise.

Re: Pdfsandwich

#30
i have a question about OCR. is there a project that lets you define fields for doing ocr, i am thinking scanning invoices and defining that this line means the invoice number, this here means the item name, item rate, etc.

many OCR software can scan this but not "understand" it to use it. ABBYY has something like this for scanning invoices so is there something for the foss folks?

Post reply on HN