Live data from Hacker News

Mistral OCR 4

mistral.ai

51–60 of 143 posts

Re: Mistral OCR 4

#51

Recently I tied OCR with Opus 4.8. (I know, not technically right tool for the job). All I needed to do was extract dates from receipts. It got about 20% of the dates wrong yet rated all as “high confidence”. Should have probably tried a more OCR specific model

> All I needed to do was extract dates from receipts

Was this... not basically a solved problem like 30 years ago? I'm pretty sure the shareware OCR tool that came with a black and white scanner I had at one point would do better than 20% wrong.

Re: Mistral OCR 4

#52

A tangential observation: the video on the linked page wasn't what I expected. I thought Mistral was a european AI company, so I didnt expect the video to be filmed in San Francisco featuring three people who don't seem to be european. I'm not against them being a global organization, that's wonderful. I was just surprised. I expected a parisian office and european accents.

~Any borderline-large European tech company will have an office on the US west coast, for sales if nothing else. And probably sales engineering. The timezone difference is eight to ten hours; there is really no way around it. (I did work for one which had an office in Vancouver, instead; same tz.)

Mistral just hired as CMO a Seattle based former Amazon/Google VP¹ , so seems their US based presence is growing.

¹ The one locally famous for being sued by Amazon for non compete back when non compete were a thing: https://www.geekwire.com/2020/amazon-sues-former-aws-marketi...

Re: Mistral OCR 4

#53

Way too expensive. Google vision OCR (which they failed to compare against), is $1.50 per 1k pages. Vs $4 from Mistral.

interesting - an equivalent Azure Document Intelligence service (scanning with layout) is 10$/1k

Re: Mistral OCR 4

#54
post #11

Do these models (this one or its competitors) do handwriting recognition?

In the sense that you can get similarity scores for individual characters referenced against a known database of characters written by various individuals. You can get stylometry scores out of small LLMs that do demographic segmentation based on writing style using the same methods. They won't have the capacity to be fed an image of handwritten text and say "Ahh, this is a note written by Winston Churchill!". You cou…

I think OP meant converting handwriting to text, not identifying a person based on their handwriting style! (but that sounds quite interesting)

Re: Mistral OCR 4

#56
It's cheap at $4/1k, but I'm hesitant to even benchmark this one again since the previous versions were all "98% accurate based on internal benchmarks of 4 pdfs" and ended up falling short of almost everything else on the market [1].

Even in this one, they just report that OlmOCRBench and OmniDocBench have "known limitations" and that's why they report flagship numbers from their internal benchmark.

https://getomni.ai/blog/benchmarking-open-source-models-for-...

Re: Mistral OCR 4

#57

A tangential observation: the video on the linked page wasn't what I expected. I thought Mistral was a european AI company, so I didnt expect the video to be filmed in San Francisco featuring three people who don't seem to be european. I'm not against them being a global organization, that's wonderful. I was just surprised. I expected a parisian office and european accents.

~Any borderline-large European tech company will have an office on the US west coast, for sales if nothing else. And probably sales engineering. The timezone difference is eight to ten hours; there is really no way around it. (I did work for one which had an office in Vancouver, instead; same tz.)

And US users spend much more than their EU counterpart
Post reply on HN