TLDR: Mistral Medium 3.5, text-only, 128B dense model, 256k context window, modified MIT license. Model is ~140G ... https://huggingface.co/mistralai/Mistral-Medium-3.5-128B They more or less claim this exceeds Claude Sonnet 3.5 on most things, but is worse than Sonnet 3.6, and exceeds all other open models. Oh and they have a cloud service that will code your apps "in the cloud". But, yeah, at this point, so does my…
Sonnet 4.5 and 4.6* There is no way it exceeds “all other” open models - but it does exceed all of Mistral’s past models. You can see it getting blown past by GLM 5.1 and Kimi in this. Still excited to give it a try
Mistral Medium 3.5
101–110 of 248 posts
Re: Mistral Medium 3.5
#102Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?
Re: Mistral Medium 3.5
#103I'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size. GLM 5.1 is an excellent model, but even at Q4 you're looking at ~400GB. Kimi K2.5 is really good too, and at Q4 quantization you're looking at almost ~600GB. This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM…
Re: Mistral Medium 3.5
#104Earlier quoted context omitted.
Qwen3.6 runs on a single GPU and beats claudes sonnet. In benchmarks and real world tests from humans. Kimi is awesome but most people won't be able to host it themselves. A lot of people are slowly realizing the moat of 1T closed source models is gone as of the last few weeks. It's going to change the industry. April was a huge month for open models, it'll be curious to see if that continues. This Mistral submission…
i run qwen 3.6. you need to drink some settle down juice.
Re: Mistral Medium 3.5
#105Re: Mistral Medium 3.5
#106Re: Mistral Medium 3.5
#107Earlier quoted context omitted.
Wow. I get that "how well can it make SVGs" isn't the (or a) gold standard for how useful a model is or isn't, but the fact the Gemma 4 26B A4B I'm running locally can blow it out of the water doesn't give me high confidence for the model. Maybe an unfair comparison, but...
It sounds like they focussed performance on not drawing svgs. Which honestly, makes a lot of sense to me.
But what stands out to me is that it's barely able to draw a "recognizable" pelican at all. The Devstral 2 model even looks slightly better, though maybe I'm splitting hairs: https://simonwillison.net/2025/Dec/9/
Re: Mistral Medium 3.5
#108Earlier quoted context omitted.
> This is the bar for Europe, huh? A few months ago China was being criticized left and right on how somehow it was not able to compete, and once DeepSeek showed up then all the hatred shifted onto how China was actually competing but exploring unfair competitive advantages. Funny how that works. Also, aren't the likes of OpenAI burning through over $2 of investment for each $1 of revenue?
[flagged]
Re: Mistral Medium 3.5
#109Compared to all other hosted LLMs that I have tested, Mistral seems to be the only one with rather strict CSP headers. When you ask them to create a website with some javascript library it will not preview, even though le chat offers canvas mode. Sometimes when a new release comes around from any provider I just want to test it a bit on the web. without paying and using an agent harness. Why are they like this ;_; Ed…
I have never wanted, needed or hoped to draw svgs with an LLM. All of the models suck at it, some are just more fun or something.
Re: Mistral Medium 3.5
#110Earlier quoted context omitted.
> China is not competing, it is distilling US models. I think you should check your notes. The likes of Kimi K2 thinking shows up as high as the second best general purpose model currently in existence. It seems they compete just fine. If you believe "distilling" is all it takes to put together a model at the top of any synthetic benchmark then I wonder what you would have to say about all US models that greatly unde…
> high as the second best general purpose model According to benchmarks which are gamed to the extreme these days. Trusting them blindly isn’t exactly rational either. They don’t necessarily translate that well to real world tasks It’s obviously not “distilling” as such but there are reasons why Chinnese models are consistently several months behind OpenAI/Antropic