OCR has been solved long time ago with vision models. Solutions are consistent, reliable, and stable. What is the point of reinventing the wheel? I would definitely understand post processing, like extracting data, answering question .. etc, but why re-doing the OCR engine itself?
Real question: what tool do you use? (for long/complex documents with tables, code, maths) - marker (with --force-ocr) gives me the best results - Mistral OCR (seems really great, but I never managed to get it work) - Mathpix (tried a long time ago) - docling (gives me garbage, I must use it wrong) - Unlimited OCR (will try it) - ???
- AWS Textract