Live data from Hacker News

Leanstral 1.5

docs.mistral.ai

21–30 of 158 posts

Re: Leanstral 1.5

#21

Real talk, does anyone use anything from Mistral because it performs the best, by whatever secular metric of your choosing? Or is it only used "because EU"? Just focus on answering the question. I wonder if anyone has observed it perform better on any objective metric in any rigorous setting.

OCR is off the charts good on every metric you can think of.

LLMs are a near-afterthought at this point if you don’t have data residency requirements. I love them and they’re slightly underrated, their models are consistently well-trained, open, but as you note, behind. There is no metric that will say they’re ahead in anything.

Re: Leanstral 1.5

#22

Real talk, does anyone use anything from Mistral because it performs the best, by whatever secular metric of your choosing? Or is it only used "because EU"? Just focus on answering the question. I wonder if anyone has observed it perform better on any objective metric in any rigorous setting.

Mistral medium is considerably better at writing than Opus. I’ve also found it very good at pulling info from pdfs. Even a complicated festival with multiple venues and timetables.

Writing what? I found it worse than gemma4 at coding even though it's 4x the parameter size

Re: Leanstral 1.5

#23
post #18
post #15

Earlier quoted context omitted.

Your reason can't be cost because there are superior models that are cheaper than Mistral models, for coding. So i re-ask the question

> Your reason can't be cost because there are superior models that are cheaper than Mistral models Nope. This is not my experience. Public pricing in token/$ is only part of the equation. Mistral tooling to consume significantly less tokens-per-given-task than the Anthropic ones. My bills currently reflects that.

I think other commenter is talking about smaller/cheaper models like Qwen that outperform mistral on just about every metric

Re: Leanstral 1.5

#24
post #18
post #15

Earlier quoted context omitted.

Your reason can't be cost because there are superior models that are cheaper than Mistral models, for coding. So i re-ask the question

> Your reason can't be cost because there are superior models that are cheaper than Mistral models Nope. This is not my experience. Public pricing in token/$ is only part of the equation. Mistral tooling to consume significantly less tokens-per-given-task than the Anthropic ones. My bills currently reflects that.

Compare to Xiaomi MiMo-V2.5 you will be shocked

Re: Leanstral 1.5

#26
What a coincidence! I just released OpenATP earlier today. OpenATP is an open-source Python package and CLI for agentic automated theorem provers. It includes support for Leanstral with Mistral’s Vibe harness. The previous production Leanstral model was deprecated on May 22nd. I will update the package to point to Leanstral 1.5 ASAP!

GitHub: https://github.com/henryrobbins/open-atp

Docs: https://open-atp.henryrobbins.com

Re: Leanstral 1.5

#27

Real talk, does anyone use anything from Mistral because it performs the best, by whatever secular metric of your choosing? Or is it only used "because EU"? Just focus on answering the question. I wonder if anyone has observed it perform better on any objective metric in any rigorous setting.

I still prefer Mistral Nemo 12B for text summarisation tasks. It has a nice style. The Mistral Small 24B is also decent. I have a YouTube transcript summariser which I like these for.

However these days I usually have Qwen 3.6 27B already loaded so I mostly just use that instead.

Re: Leanstral 1.5

#28

Interesting that this only specialized for Lean4 and not for similar like Coq

I would have preferred actual proof objects, as in Metamath's: separate the actual proof from the heuristics used to find it (also valuable, but a different thing).

Re: Leanstral 1.5

#29

Real talk, does anyone use anything from Mistral because it performs the best, by whatever secular metric of your choosing? Or is it only used "because EU"? Just focus on answering the question. I wonder if anyone has observed it perform better on any objective metric in any rigorous setting.

I liked that their website didn’t ask for my phone number, IIRC.
Post reply on HN