I want to believe it's gonna be good, but after trying GPT-5.5 even the most advanced Chinese models seem depressing.
Mistral Medium 3.5
11–20 of 248 posts
Re: Mistral Medium 3.5
#12Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?
Re: Mistral Medium 3.5
#13It's okay, nothing exceptional, but any news from non US and non Chinese models is still good news.
Re: Mistral Medium 3.5
#14Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?
If it's not US and it's within a few percent of SOTA that might be good enough for a lot of people (eg Europeans)
Re: Mistral Medium 3.5
#15Earlier quoted context omitted.
Then you’ll be happy to learn it’s not Chinese
GP is stating that the second best in the field, the Chinese, is so far behind the best in the field, GPT 5.5, that it is not even worth testing anything else.
Re: Mistral Medium 3.5
#16Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?
Re: Mistral Medium 3.5
#17Re: Mistral Medium 3.5
#18Pre-agent, there wasn't always an obvious difference between models. Various models had their charms. Nowadays, I don't want to entertain anything less than the frontier models. The difference in capability is enormous and choosing anything less has a real cost in terms of productivity.
I've been a big fan of the smaller labs like Mistral and especially Cohere but it's been a while since I've been excited by a release by either company.
That said, I'm using mistral voxtral realtime daily – it's great.
Re: Mistral Medium 3.5
#19The different results on some benchmarks vibes as if this is truly an independently trained model, not just exfiltrated frontier logs, which I think is also really important - having different weight architectures inside a particular model seems like a benefit on its own when viewed from a global systems architecture perspective.
Re: Mistral Medium 3.5
#20One thing in particular I was disappointed in was its bad explanations when asking about French grammar. It made multiple mistakes and the other models got it right, even Qwen 3.6 27b!
Anyway, I'm hoping they catch up some more.