Leo, i benchmarked Qwen2.5-VL-3B, 4-bit via MLX, against Tesseract on the same 24 samples: https://thiagotigaz.github.io/ocr-it/bench/

Its much lower and the error rate is much higher. It works, but for our usecase, "clean rendered text" (extract from kindle for example) tesseract is much better.