Live data from Hacker News

Mistral Medium 3.5

mistral.ai

11–20 of 248 posts

Re: Mistral Medium 3.5

#12

Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?

If it's not US and it's within a few percent of SOTA that might be good enough for a lot of people (eg Europeans)

Re: Mistral Medium 3.5

#14
post #12

Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?

If it's not US and it's within a few percent of SOTA that might be good enough for a lot of people (eg Europeans)

Gemma has been better for us at EU languages than mistral (for comparable sized models) :/ so ... dunno. What mistral does well and others are lagging behind is deploying on prem with their engineers and know-how, offering tuned models for your tasks and finetuning on your own data. (I expect google to start offering this next)

Re: Mistral Medium 3.5

#15
post #3

Earlier quoted context omitted.

Then you’ll be happy to learn it’s not Chinese

GP is stating that the second best in the field, the Chinese, is so far behind the best in the field, GPT 5.5, that it is not even worth testing anything else.

Thanks for the translation, I did not express it very clearly. Anything that I try is so much worse.

Re: Mistral Medium 3.5

#16

Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?

Price and speed.

Re: Mistral Medium 3.5

#17
I'm rooting for Mistral. It seems they are making a big bet that smaller models will win over larger ones and I can see it happening. I was running some simple chat and tool-calling benchmarks for small models and Mistral Small 4 performed well for it's price ($.15/$.60). Seeing this today got me excited, benchmarks seems solid compared to models much larger, but it's priced higher than Haiku, 5.4 mini, and all the the Chinese models it's comparing itself too. It's not even winning those benches either, just being competitive with them, which is great, those models are 5x+ the size, but they are also 1/2 the price. Hard to be excited about that.

Re: Mistral Medium 3.5

#18
This release Mistral really reminds you of the gap between the frontier labs and everyone else.

Pre-agent, there wasn't always an obvious difference between models. Various models had their charms. Nowadays, I don't want to entertain anything less than the frontier models. The difference in capability is enormous and choosing anything less has a real cost in terms of productivity.

I've been a big fan of the smaller labs like Mistral and especially Cohere but it's been a while since I've been excited by a release by either company.

That said, I'm using mistral voxtral realtime daily – it's great.

Re: Mistral Medium 3.5

#19
As always, rooting for these guys — model and national diversity is great. This looks like a solid foundation to build on; hopefully the 3.6/3.7 will dial in more gains. It looks like maybe from the computer use benchmarks that their vision pipeline could use improvement, but that’s just speculation.

The different results on some benchmarks vibes as if this is truly an independently trained model, not just exfiltrated frontier logs, which I think is also really important - having different weight architectures inside a particular model seems like a benefit on its own when viewed from a global systems architecture perspective.

Re: Mistral Medium 3.5

#20
I use Mistral Le Chat quite a bit.

One thing in particular I was disappointed in was its bad explanations when asking about French grammar. It made multiple mistakes and the other models got it right, even Qwen 3.6 27b!

Anyway, I'm hoping they catch up some more.

Post reply on HN