Live data from Hacker News

Mistral Medium 3.5

mistral.ai

101–110 of 248 posts

Re: Mistral Medium 3.5

#101
post #27
post #6

TLDR: Mistral Medium 3.5, text-only, 128B dense model, 256k context window, modified MIT license. Model is ~140G ... https://huggingface.co/mistralai/Mistral-Medium-3.5-128B They more or less claim this exceeds Claude Sonnet 3.5 on most things, but is worse than Sonnet 3.6, and exceeds all other open models. Oh and they have a cloud service that will code your apps "in the cloud". But, yeah, at this point, so does my…

Sonnet 4.5 and 4.6* There is no way it exceeds “all other” open models - but it does exceed all of Mistral’s past models. You can see it getting blown past by GLM 5.1 and Kimi in this. Still excited to give it a try

It looks like qwen 3.6 is winning and smaller for the April small model roll out

Re: Mistral Medium 3.5

#102

Looks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?

They did not stop using it due to contamination. They said it's flawed and indirectly said anthropics results were impossible. It's very possible they are sore losers

Re: Mistral Medium 3.5

#103
post #52

I'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size. GLM 5.1 is an excellent model, but even at Q4 you're looking at ~400GB. Kimi K2.5 is really good too, and at Q4 quantization you're looking at almost ~600GB. This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM…

The competition is on DeepSeek v4 Flash for similar size / deployment target.

Re: Mistral Medium 3.5

#104

Earlier quoted context omitted.

Qwen3.6 runs on a single GPU and beats claudes sonnet. In benchmarks and real world tests from humans. Kimi is awesome but most people won't be able to host it themselves. A lot of people are slowly realizing the moat of 1T closed source models is gone as of the last few weeks. It's going to change the industry. April was a huge month for open models, it'll be curious to see if that continues. This Mistral submission…

i run qwen 3.6. you need to drink some settle down juice.

No way it's awesome.

Re: Mistral Medium 3.5

#106
With most OSS releases being MoEs, and modern GPUs optimized for MoEs, can somebody with knowledge of the topic explain or speculate why Mistral might have opted for a dense model?

Re: Mistral Medium 3.5

#107
post #69

Earlier quoted context omitted.

Wow. I get that "how well can it make SVGs" isn't the (or a) gold standard for how useful a model is or isn't, but the fact the Gemma 4 26B A4B I'm running locally can blow it out of the water doesn't give me high confidence for the model. Maybe an unfair comparison, but...

It sounds like they focussed performance on not drawing svgs. Which honestly, makes a lot of sense to me.

Drawing SVGs isn't something I really care about either, and I think it's still to "qualitatively compare" e.g. "Opus's pelican vs GPT's pelican vs GLM's pelican" or whatever the kids are doing.

But what stands out to me is that it's barely able to draw a "recognizable" pelican at all. The Devstral 2 model even looks slightly better, though maybe I'm splitting hairs: https://simonwillison.net/2025/Dec/9/

Re: Mistral Medium 3.5

#108
post #35

Earlier quoted context omitted.

> This is the bar for Europe, huh? A few months ago China was being criticized left and right on how somehow it was not able to compete, and once DeepSeek showed up then all the hatred shifted onto how China was actually competing but exploring unfair competitive advantages. Funny how that works. Also, aren't the likes of OpenAI burning through over $2 of investment for each $1 of revenue?

[flagged]

I find it funny how people don't realize the technical achievements and papers coming out of deepseek or Alibaba. They are making this whole AI thing sustainable and cheap and available to do at home. That's the future. I should be able to run my own harness and model and never bother with openai or anthropic at all.

Re: Mistral Medium 3.5

#109
post #61

Compared to all other hosted LLMs that I have tested, Mistral seems to be the only one with rather strict CSP headers. When you ask them to create a website with some javascript library it will not preview, even though le chat offers canvas mode. Sometimes when a new release comes around from any provider I just want to test it a bit on the web. without paying and using an agent harness. Why are they like this ;_; Ed…

I have never wanted, needed or hoped to draw svgs with an LLM. All of the models suck at it, some are just more fun or something.

I can't speak for what you consider sucking, but there is a significant difference between Mistral and Kimi or Gemini. I find the others to be usable for my needs.

Re: Mistral Medium 3.5

#110
post #91

Earlier quoted context omitted.

> China is not competing, it is distilling US models. I think you should check your notes. The likes of Kimi K2 thinking shows up as high as the second best general purpose model currently in existence. It seems they compete just fine. If you believe "distilling" is all it takes to put together a model at the top of any synthetic benchmark then I wonder what you would have to say about all US models that greatly unde…

> high as the second best general purpose model According to benchmarks which are gamed to the extreme these days. Trusting them blindly isn’t exactly rational either. They don’t necessarily translate that well to real world tasks It’s obviously not “distilling” as such but there are reasons why Chinnese models are consistently several months behind OpenAI/Antropic

[dead]
Post reply on HN