I find this utterly bizarre. Once upon a time, if you wanted to left pad a string, you would just do it. A while later, people discovered that you could use a library. (I’m joking a bit here, but libraries are genuinely useful.). With a library, you get to pick from various schemes and schedules for updating the library, but you have a degree of control. But now apparently you’re supposed to use a web API and depend…
This really does not resonate at all, and I have the scars to prove it. I used to work on a browser-based document management system, and I would have used (or at least tried) all of these APIs without hesitation. PDFs are a pain and the mish mash of poor functioning tools that exist provides a constant headache. 1) OCR'ing of a PDF is difficult. The only good service is Google, but requires that you break it into pa…
OCRspace is OK, too, and easier to use. You can just send the PDF. It is free for PDFs with 3 or less pages.
> 2) People want searchable OCR'd PDFs where you can highlight the text, even when it's a bitmap underneath.
OCRspace can also create searchable PDFs: https://ocr.space/searchablepdf