Live data from Hacker News

Mistral Medium 3.5

mistral.ai

21–30 of 248 posts

Re: Mistral Medium 3.5

#21
post #2

I want to believe it's gonna be good, but after trying GPT-5.5 even the most advanced Chinese models seem depressing.

I am not following this obsession with SOTA and benchmark rankings

I have been using DeepSeek and GLMnmodels with OpenCode and Codex and Claudr side by side.

I have not found the Chinese models lacking. I enjoy for coding and like to maintain full control of my codebade and deeply care about the GOF patterns. So I am very stringent in terms of what I want the LLM to code and how to code.

So from my perspective, they are all about the same.

Re: Mistral Medium 3.5

#23
post #20

I use Mistral Le Chat quite a bit. One thing in particular I was disappointed in was its bad explanations when asking about French grammar. It made multiple mistakes and the other models got it right, even Qwen 3.6 27b! Anyway, I'm hoping they catch up some more.

There's a good chance that they'll catch up. The "AI race" is a race to the bottom, with the leaders blowing huge wads of cash on capabilities that get replicated months later by the competition at a fraction of the cost.

The only benefit of leading is mindshare. OpenAI is doubling down on that, by investing in communication companies. That's their pathetic attempt at a "moat".

Re: Mistral Medium 3.5

#24
post #2

I want to believe it's gonna be good, but after trying GPT-5.5 even the most advanced Chinese models seem depressing.

I am not following this obsession with SOTA and benchmark rankings I have been using DeepSeek and GLMnmodels with OpenCode and Codex and Claudr side by side. I have not found the Chinese models lacking. I enjoy for coding and like to maintain full control of my codebade and deeply care about the GOF patterns. So I am very stringent in terms of what I want the LLM to code and how to code. So from my perspective, they…

That I agree with, but for more complex autonomous changes the differences are considerable. However, it seems that most models will reach the saturation time in which they will be useful for almost everything and the difference will be in more and more niche and specialized tasks.

Re: Mistral Medium 3.5

#25
post #3

Earlier quoted context omitted.

Then you’ll be happy to learn it’s not Chinese

GP is stating that the second best in the field, the Chinese, is so far behind the best in the field, GPT 5.5, that it is not even worth testing anything else.

Is GPT 5.5 the best in the field? I think Opus is still better despite Anthropic's recent stumbling.

Re: Mistral Medium 3.5

#26
post #6

TLDR: Mistral Medium 3.5, text-only, 128B dense model, 256k context window, modified MIT license. Model is ~140G ... https://huggingface.co/mistralai/Mistral-Medium-3.5-128B They more or less claim this exceeds Claude Sonnet 3.5 on most things, but is worse than Sonnet 3.6, and exceeds all other open models. Oh and they have a cloud service that will code your apps "in the cloud". But, yeah, at this point, so does my…

You mean Sonnet 4.5 and 4.6 riight

Re: Mistral Medium 3.5

#27
post #6

TLDR: Mistral Medium 3.5, text-only, 128B dense model, 256k context window, modified MIT license. Model is ~140G ... https://huggingface.co/mistralai/Mistral-Medium-3.5-128B They more or less claim this exceeds Claude Sonnet 3.5 on most things, but is worse than Sonnet 3.6, and exceeds all other open models. Oh and they have a cloud service that will code your apps "in the cloud". But, yeah, at this point, so does my…

Sonnet 4.5 and 4.6*

There is no way it exceeds “all other” open models - but it does exceed all of Mistral’s past models.

You can see it getting blown past by GLM 5.1 and Kimi in this.

Still excited to give it a try

Re: Mistral Medium 3.5

#28
I like the idea of Mistral, but the last time I evaluated Mistral Vibe it was really nice for $15/month but not as effective as Gemini Plus with AntiGravity and gemini-cli. I am currently running Gemini Ultra on a 3 month 'special deal' and AntiGravity with Opus 4.7 tokens is pretty much fantastic.

That said, when I stop spending money on Gemini Ultra, I will give Mistral Vibe another 1-month test.

I like the entire business model and vibe of Mistral so much more than OpenAI/Anthropic/Google but I also have stuff to get done. I am curious if Mistral Vibe for $15/month is a stable business model (i.e., can they make a profit).

Re: Mistral Medium 3.5

#29

This release Mistral really reminds you of the gap between the frontier labs and everyone else. Pre-agent, there wasn't always an obvious difference between models. Various models had their charms. Nowadays, I don't want to entertain anything less than the frontier models. The difference in capability is enormous and choosing anything less has a real cost in terms of productivity. I've been a big fan of the smaller l…

> Pre-agent, there wasn't always an obvious difference between models. Various models had their charms. Nowadays, I don't want to entertain anything less than the frontier models. The difference in capability is enormous and choosing anything less has a real cost in terms of productivity.

It's just apples to oranges.

There is not a clear, across the board, winner on non-agentic tasks between Gemini, ChatGPT, and Claude - the simple chatbot interface.

But Claude Code is substantially better than Codex which itself is notably better than Gemini-cli.

In this vein, it should not be surprising that Claude Code is way better than non-frontier models for agentic coding... It's substantially better than other frontier models at specialized agentic tasks.

Re: Mistral Medium 3.5

#30
post #6

TLDR: Mistral Medium 3.5, text-only, 128B dense model, 256k context window, modified MIT license. Model is ~140G ... https://huggingface.co/mistralai/Mistral-Medium-3.5-128B They more or less claim this exceeds Claude Sonnet 3.5 on most things, but is worse than Sonnet 3.6, and exceeds all other open models. Oh and they have a cloud service that will code your apps "in the cloud". But, yeah, at this point, so does my…

Unfortunately they only compare to old “all other open models”. There are probably over 10 other open models better than it by now.
Post reply on HN