Live data from Hacker News

Mistral OCR 4.1

docs.mistral.ai

91–100 of 182 posts

Re: Mistral OCR 4.1

#91

I've got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc. Nothing special about this model for overly-detailed work like mine. It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could p…

While I haven't tried OpenAI for OCR, I've put my small scale OCR work through both Claude and Mistral OCR. Claude is absolutely better - even in OCR work I did last week and compared with Mistral OCR 4.0. Mistral's one advantage is that Anthropic now flags OCR, because they don't allow anything that could be considered "reproduction", even of work for which you own the copyright. So my new workflow is Mistral OCR fo…

>Anthropic now flags OCR

I haven't seen any difference in my ocr workflows, what do you mean by this?

Re: Mistral OCR 4.1

#92

I've got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc. Nothing special about this model for overly-detailed work like mine. It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could p…

I made a benchmark for handwriting recognition for a project while keeping line breaks and errors (grammar, spelling). Sonnet absolutely dominates it since a good half a year. 5.6 did not change that for me. This should also translate to better ocr.

Re: Mistral OCR 4.1

#93

I've got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc. Nothing special about this model for overly-detailed work like mine. It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could p…

While I haven't tried OpenAI for OCR, I've put my small scale OCR work through both Claude and Mistral OCR. Claude is absolutely better - even in OCR work I did last week and compared with Mistral OCR 4.0. Mistral's one advantage is that Anthropic now flags OCR, because they don't allow anything that could be considered "reproduction", even of work for which you own the copyright. So my new workflow is Mistral OCR fo…

> Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check.

What does this entail? What does Claude do to decide that the text it was provided was hallucinated? Are you telling Claude that the source was OCR'd by another LLM?

Re: Mistral OCR 4.1

#94

The VLM's are so good at complex document understanding now. But you just can't trust them not to invisibly censor sensitive clinical/legal docs, even at the maximally permissive settings. And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex docum…

The way my harness set it up is going through 2 or 3 providers, and cross-checking through them, also with plain text extracted if available.

I think we also had a layer that for any quote extracted tested it back if it exists within the original.

If you wanted 100% accuracy, I think it wouldn't be too difficult nowadays to guess the font&size&other text settings, and re render the crucial parts.

Re: Mistral OCR 4.1

#95

At this point I lost all hope for Europe playing any significant role in the AI race. If that’s a good or a bad thing I don’t know, but it seems to me like that’s the reality.

And US lost the significant role in chip/pc manufacturing. But if a product becomes commodity or utility (which at least for now it seems is the direction), with little lockin, it's not a big deal.

I hope we (EU) don't waste money trying to train local models (which at least some people in Poland try to do), and tries to build our own chips - AI chips have different architecture than regular processor/GPU, and TSMC doesn't need to be winner in this new race.

And if not this, then smaller labs, harnesses and actual application.

Re: Mistral OCR 4.1

#96

Earlier quoted context omitted.

While I haven't tried OpenAI for OCR, I've put my small scale OCR work through both Claude and Mistral OCR. Claude is absolutely better - even in OCR work I did last week and compared with Mistral OCR 4.0. Mistral's one advantage is that Anthropic now flags OCR, because they don't allow anything that could be considered "reproduction", even of work for which you own the copyright. So my new workflow is Mistral OCR fo…

> Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check. What does this entail? What does Claude do to decide that the text it was provided was hallucinated? Are you telling Claude that the source was OCR'd by another LLM?

I assume they pass the image alongside the text to redo/recheck the work. Yes this is silly, but is apparently required to get around refusals.

Re: Mistral OCR 4.1

#97
post #56

Earlier quoted context omitted.

Citing nuclear weapons isn't the flex you think it is. There are only 9 countries that have nuclear weapons and they absolutely flex this power over non-nuclear powers (see Ukraine, Germany, SE Asia etc..) Europe already has an innovation problem that's already causing structural economic instabilities which Germany has been struggling (and lately failing) to prop up. As much as it pains me to say this, AI is already…

> There are only 9 countries that have nuclear weapons QED being first didn't grant exclusivity. The premise wasn't that AGI isn't useful. There are actually layers to the metaphor where first-mover advantage of AGI is even less meaningful than it was for nuclear weapons, but I leave those as an exercise to the reader to discover.

Nobody argued exclusivity. But there have been clear advantages to having nukes first.

Re: Mistral OCR 4.1

#98
post #92

I've got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc. Nothing special about this model for overly-detailed work like mine. It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could p…

I made a benchmark for handwriting recognition for a project while keeping line breaks and errors (grammar, spelling). Sonnet absolutely dominates it since a good half a year. 5.6 did not change that for me. This should also translate to better ocr.

You're the second person to mention handwriting. I think it might be handwriting-specific; perhaps Anthropic has a better corpus for this.

I really wouldn't know, though. Anthropic models barf out copyright issues for my use case, so I'm unable even to benchmark them. It's a common problem when you're scanning public domain books. Mine are reference texts often cited.

Re: Mistral OCR 4.1

#99
post #58

At this point I lost all hope for Europe playing any significant role in the AI race. If that’s a good or a bad thing I don’t know, but it seems to me like that’s the reality.

Much like spaceflight, aerospace or nuclear engineering, you need to retain local talent for national defense purposes. Being 70% as good is still way better than being 100% dependent and heavily leveraged by your opponents.

Being 70% as good in nuclear engineering sounds scary AF. How would you rank Chernobyl? Better or worse than 70%?

Re: Mistral OCR 4.1

#100

The VLM's are so good at complex document understanding now. But you just can't trust them not to invisibly censor sensitive clinical/legal docs, even at the maximally permissive settings. And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex docum…

> But you just can't trust them not to invisibly censor sensitive clinical/legal docs

What's an example of this?

Post reply on HN