Live data from Hacker News

Mistral OCR 3

mistral.ai

31–40 of 137 posts

Re: Mistral OCR 3

#31
post #20

Gave it a birth registry from a Portuguese locality from 1755 which my dad and I often decipher to figure out geneology and it did a terrible job. Regular Gemini Thinking can actually get 70-80% of the documents correct except lots of mistakes on given names. Chatgpt maybe understands like 50-60%. This Mistral model butchered the whole text, literally not a word was usable. To the point I think I'm doing something wr…

Oh god, I'm sure I wouldn't come close to 50%; that's so hard to read

It's tough but my dad is quite good at it. He has books of common abbreviations and agglutinations from different centuries. After you get used to it it's faster and very fun.

We were mind blown how good Gemini was at it.

Re: Mistral OCR 3

#32
post #31

Earlier quoted context omitted.

Oh god, I'm sure I wouldn't come close to 50%; that's so hard to read

It's tough but my dad is quite good at it. He has books of common abbreviations and agglutinations from different centuries. After you get used to it it's faster and very fun. We were mind blown how good Gemini was at it.

I am too. Gemini 3.0 fast on old scrawled diary entries in English from 100+ years ago got them 95% right. It also added historical context when I prefaced the images with the identity of the writer, such as summaries of an old military unit history in Europe post-WW1 it got from a very obscure U.S. Army archive.

Huge timesaver.

Re: Mistral OCR 3

#33
post #11
post #2

there has been so many open source OCR in the last 3 months that would be good to compare to those especially when some are not even 1B params and can be run on edge devices. - paddleOCR-VL - olmOCR-2 - chandra - dots.ocr I kind of miss there is not many leaderboard sections or arena for OCR and CV and providers hosting those. Neglected on both Artificial Analysis and OpenRouter.

what I like in MistralOCR is that they have simple pricing $1/1k pages and API hosted on their servers. With other OCR is hard to compare pricing because are token based and you don't know how many tokens is the image unless you run your own test. E.g. with Gemini 3.0 flash you might seem that model pricing increased only slightly comparing to Gemini 2.5 flash until you test it and will see that what used to be 258 p…

But they doubled the price g for this new mistralocr3 model to 2$

Re: Mistral OCR 3

#34
Does it handle math expressions (those rendered from LaTeX) well? I've been looking for a good OCR model to transcribe my math textbooks into markdown (obviously ignoring the images and figures) with LaTeX as math expressions, and none of the current OCR models work reliably enough.

EDIT: you can try it yourself for free at https://console.mistral.ai/build/document-ai/ocr-playground once you create a developer account! Fingers crossed to see how well it works for my use case.

Re: Mistral OCR 3

#35
post #29
post #18

Earlier quoted context omitted.

Someone posted a project here about a month ago where they compare models in head-to-head matchups similar to llmarena https://www.ocrarena.ai/leaderboard Hasn't been updated for Mistral but so far gemeni seems to top the leaderboard.

How can something have a very high ELO but a very low win rate?

You don't loose any elo if your opponent is much stronger than you. Remis could in theory play a part as well.

Re: Mistral OCR 3

#36
post #3

From a tweet: https://x.com/i/status/2001821298109120856 > can someone help folks at Mistral find more weak baselines to add here? since they can't stomach comparing with SoTA.... > (in case y'all wanna fix it: Chandra, dots.ocr, olmOCR, MinerU, Monkey OCR, and PaddleOCR are a good start)

I'd want to see a comparison with Qwen 3 VL 235B-A22B, which is IME significantly better than MinerU.

Re: Mistral OCR 3

#37
post #2

there has been so many open source OCR in the last 3 months that would be good to compare to those especially when some are not even 1B params and can be run on edge devices. - paddleOCR-VL - olmOCR-2 - chandra - dots.ocr I kind of miss there is not many leaderboard sections or arena for OCR and CV and providers hosting those. Neglected on both Artificial Analysis and OpenRouter.

[dead]

Re: Mistral OCR 3

#38
post #20

Gave it a birth registry from a Portuguese locality from 1755 which my dad and I often decipher to figure out geneology and it did a terrible job. Regular Gemini Thinking can actually get 70-80% of the documents correct except lots of mistakes on given names. Chatgpt maybe understands like 50-60%. This Mistral model butchered the whole text, literally not a word was usable. To the point I think I'm doing something wr…

Just gave it a shot with Grok 4.1 thinking - do you have the ground truth translation to compare? I've tried 4 different times, with slight tweaks adding information from your description, and it's given me a range of interpretations. It'd be nice to see if any of them got close - a couple were more like pulpy telenovela plots, lol.

The model might need tuning in order to be effective - this is normal for releases of image mode models, and after a couple days, there will be properly set up endpoints to test from, so it might be much better than you think. Or it could be really bad with turn of the 19th century portugese cursive.

Post reply on HN