Live data from Hacker News

Mistral OCR 4

mistral.ai

101–110 of 143 posts

Re: Mistral OCR 4

#101
post #87
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

The USPS Remote Encoding Center in Salt Lake City examined 841,260,847 images of poorly written addresses in fiscal year 2025. [0] Unfortunately the page does not have a base rate--the total number of mail pieces that were not prepared for automated processing. Total first class mail, which includes a lot of bills prepared for automation was 25.7 billion [1]. If 10% of that are non-automated, then .8 / 2.57 = .31 or…

I can't help you with any of those questions... but back in the 90s I use to be one of those employees that looked at the image on the screen and typed the address information in Salt Lake City.

Quantitatively, I don't know the stats, but qualitatively I can confirm it felt like a lot.

Re: Mistral OCR 4

#102
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

USPS can also make a list of every address, that's not possible generally with freeform text

Re: Mistral OCR 4

#103

Earlier quoted context omitted.

I do not believe this story. Opus 4.8 scanned hundreds of PDFs for me recently with the worst handwriting imaginable. 100% successful, other than one record where even I could not figure out what was written.

I believe it. Makes me curious what your prompt was that got such a good result out of Opus.

likely just opus being dumbed down to prepare for fable

Re: Mistral OCR 4

#104

Earlier quoted context omitted.

Unfortunately Europeans are terrible customers for making money. They ask a lot of questions and they're very stingy with their wallets. Americans on the other hand ...

This is absolutely not why there are no leading AI, other important silicon tech, or relevant space companies in Europe. To some degree they exist but are all B-Tier in comparison to US/China. You'd be surprised just how lose money can sit in Europe, I guess. Just not the way it needs to be for this. The financial structure of the EU is nowhere close to enabling these capital devouring endeavors based on lofty future…

[flagged]

Re: Mistral OCR 4

#105
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

My father once received a letter from Algeria, with 3 words on the envelope : his first name, "Créteil" (the town where he lived, ≈100k inhabitants), and "France". Of course, in the 70s there was no Internet nor central database to find him, yet the postal service managed to deliver the letter. He was a very active social worker, managed a youth football team, etc. which made him locally well-known by his first name.

Nowadays, many people can't find anyone or any place unless their phone helps them. And postmen never stop to chat. Such a letter would not pass through the technology process, and probably not through the human network.

Re: Mistral OCR 4

#106
> On our internal multilingual evaluation, OCR 4 leads across all eight language groups — English, Western Europe, Eastern Europe, Middle Eastern, Chinese, East Asian, Southeast Asian, and specialized languages (Hindi, Japanese, Georgian, Bengali, Armenian, Hebrew, Greek, Gujarati, Tamil, Malayalam, Kannada, Telugu).

The initial version of this page called these "minor languages" (vs specialized language), which is telling. If you're a speaker of one of these: This is why you need a sovereign set of models. (Japanese government: Are you listening?)

Re: Mistral OCR 4

#107
post #94

Earlier quoted context omitted.

To the best of my knowledge, most of the founding team started their careers in the US ( meta,etc..) and their primary investors are US VCs. In that regard, they smartly benefit on both side : US funding and European brains

Uhm... isnt mistral mostly funded by ASML? A dutch company?

No, it's just one of the investors/customers.

Re: Mistral OCR 4

#108

After paying for Mistral and using it for a while I genuinely hated it. It's a productivity black hole and can't realistically compete with anyone. I chose it only because it was European, but no. I'd rather let my one year subscription go to waste than use anything 'Mistral'.

Codestral is pretty good for coding autocomplete. possibly better than even Cursor's

Re: Mistral OCR 4

#109
post #87
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

The USPS Remote Encoding Center in Salt Lake City examined 841,260,847 images of poorly written addresses in fiscal year 2025. [0] Unfortunately the page does not have a base rate--the total number of mail pieces that were not prepared for automated processing. Total first class mail, which includes a lot of bills prepared for automation was 25.7 billion [1]. If 10% of that are non-automated, then .8 / 2.57 = .31 or…

I’m not totally sure what your point is, but my response is that most OCR technology is reading “automated” (i.e. computer-printed) documents such as PDFs and things like that. So I think parsing the numbers by “automated” vs “non-automated” is not a very helpful way to think about the success of USPS OCR technology; the gross percentage of manual reviews compared to total mail volume is a much better way at looking at the success of their OCR. That’s my perspective anyway, but maybe commercial OCR is really optimized for reading handwriting and I’m just not aware of it. I’m not an expert in the area.

Re: Mistral OCR 4

#110
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

>... US addresses are absurdly non-standardized.

Laughs in Indian addresses.

Post reply on HN