Just tested with a multilingual (bidi) English/Hebrew document. The Hebrew output had no correspondence to the text whatsoever (in context, there was an English translation, and the Hebrew produced was a back-translation of that). Their benchmark results are impressive, don't get me wrong. But I'm a little disappointed. I often read multilingual document scans in the humanities. Multilingual (and esp. bidi) OCR is ch…
You can get bounding boxes from our pdf api at Mathpix.com Disclaimer, I’m the founder
There are a few annoying issues, but overall I am very happy with it.