Live data from Hacker News

Mistral OCR 4

mistral.ai

111–120 of 143 posts

Re: Mistral OCR 4

#112
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

My father once received a letter from Algeria, with 3 words on the envelope : his first name, "Créteil" (the town where he lived, ≈100k inhabitants), and "France". Of course, in the 70s there was no Internet nor central database to find him, yet the postal service managed to deliver the letter. He was a very active social worker, managed a youth football team, etc. which made him locally well-known by his first name.…

> And postmen never stop to chat

I can say that at least where I live (Brisbane Australia's inner-ring suburbs) that's still the case: my postie is super friendly, loves my dog, and always has a couple of minutes to say hi!

And its great, because I live at a house number that is A, and there is a ... but no B or anything, which trips up a surprising amount of people/delivery drivers, so having a postie who cares helps

Re: Mistral OCR 4

#113
post #109
post #87

Earlier quoted context omitted.

The USPS Remote Encoding Center in Salt Lake City examined 841,260,847 images of poorly written addresses in fiscal year 2025. [0] Unfortunately the page does not have a base rate--the total number of mail pieces that were not prepared for automated processing. Total first class mail, which includes a lot of bills prepared for automation was 25.7 billion [1]. If 10% of that are non-automated, then .8 / 2.57 = .31 or…

I’m not totally sure what your point is, but my response is that most OCR technology is reading “automated” (i.e. computer-printed) documents such as PDFs and things like that. So I think parsing the numbers by “automated” vs “non-automated” is not a very helpful way to think about the success of USPS OCR technology; the gross percentage of manual reviews compared to total mail volume is a much better way at looking…

My response was to your thesis: "whenever I see announcements about OCR it feels like this should be a solved problem if it’s been accomplished at the scale of USPS for many years." I think the USPS has "solved" much of the problem by getting the most prolific generators of mail to conform with more basic tech than OCR, the bar code.

Much commercial mail (including first class non-junk mail) is physically presorted and bundled as it is dropped with USPS and has a bar code that states the routing needed. Stuff that has had OCR performed by computer or human gets a little sticker near the bottom with the barcode.

The barcode is applied by the sender; the Postal Service required use of the Intelligent Mail barcode to qualify for automation prices beginning January 28, 2013. Use of the barcode provides increased overall efficiency, including improved deliverability, and new services.

https://en.wikipedia.org/wiki/Intelligent_Mail_barcode

Re: Mistral OCR 4

#114
post #88

A tangential observation: the video on the linked page wasn't what I expected. I thought Mistral was a european AI company, so I didnt expect the video to be filmed in San Francisco featuring three people who don't seem to be european. I'm not against them being a global organization, that's wonderful. I was just surprised. I expected a parisian office and european accents.

Another company like this is Blackmagic Design. Despite being overwhelmingly based in Australia, you'd think it was an American company based on office listing ordering on https://www.blackmagicdesign.com/company/offices and /company page.

...I had no idea Blackmagic was australian! Thats wild, but maybe explains why their tech is so prevalent here haha. Well that and its all pretty excellent and great value (at least in the indie scene I was a part of)

Re: Mistral OCR 4

#115
post #77

I’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge…

There's a lot of weird edge cases with US addresses. Carmel by the sea doesn't have street numbers. Florida keys addresses are often just a mile marker. The mail gets delivered because a human on the route is familiar with them.

Re: Mistral OCR 4

#116

Earlier quoted context omitted.

Unfortunately Europeans are terrible customers for making money. They ask a lot of questions and they're very stingy with their wallets. Americans on the other hand ...

This is absolutely not why there are no leading AI, other important silicon tech, or relevant space companies in Europe. To some degree they exist but are all B-Tier in comparison to US/China. You'd be surprised just how lose money can sit in Europe, I guess. Just not the way it needs to be for this. The financial structure of the EU is nowhere close to enabling these capital devouring endeavors based on lofty future…

I think “Just not the way it needs to be for this.” is exactly the point.

Re: Mistral OCR 4

#117
Are there any open models focused on LPR (license plate recognition)?

I have found some old ones but curious if there are new ones being developed like this OCR model. I may even try it for the purpose and see if it does well.

Re: Mistral OCR 4

#118

Earlier quoted context omitted.

Why would anybody do that you would simply get terrible results compared to dozens of other more capable models. It's for converting to text not answering questions. Just seems like you need some sort of weird angle to bring out an anti AI stance

I think his comment is referring to a scenario where a decision is made on financial numbers that are misrecognized. E.g. 9.0% actual is OCR’d as 90%

I don't think so

But anyways just a side note one way to help reduce these errors is if you pass in both the original image and the OCR'd text to the models that make the decisions

Re: Mistral OCR 4

#120
post #75

All AI labs really need to stop using truncated y-axes for benchmark bar charts... https://mistral.ai/_astro/cm-engish_ZhlvoT.webp?dpl=6a3a94bd...

No, they need to keep using truncated y-axes to increase the hype cycle.
Post reply on HN