Pelicans are OK but not earth-shattering: https://simonwillison.net/2025/Dec/2/introducing-mistral-3/
Mistral 3 family of models released
121–130 of 243 posts
Re: Mistral 3 family of models released
#122I was subscribing to these guys purely to support the EU tech scene. So I was on Pro for about 2 years while using ChatGPT and Claude. Went to actually use it, got a message saying that I missed a payment 8 months previously and thus wasn't allowed to use Pro despite having paid for Pro for the previous 8 months. The lady I contacted in support simply told me to pay the outstanding balance. You would think if you mis…
This seems like a legitimate complaint... I wonder why it's downvoted
Also a lot of Europeans are upset at US tech dominance. It's a position we've roped ourselves in to so any commentary that criticises an EU tech success story is seen as being unnecessarily negative.
However I do mean it as a warning to others, I got burned even with good intentions.
Re: Mistral 3 family of models released
#123Earlier quoted context omitted.
Hard to gauge what gibberish is without an example of the data and what you prompted the LLM with.
If you wanted examples, you needed only ask :) These are screenshots from that week: https://x.com/barrelltech/status/1995900100174880806 I'm not going to share the prompt because (1) it's very long (2) there were dozens of variations and (3) it seems like poor business practices to share the most indefensible part of your business online XD
Impressive, I haven't seen that myself yet, I've only used 5 conversationally, not via API yet.
Re: Mistral 3 family of models released
#124Earlier quoted context omitted.
Are they ahead of all other recent open models? Is there a leaderboard?
There is a leaderboard [1] but we'll have to wait till april for the competition to end to know what models they're using. The current number 3 on there (34/50) has mentioned in discussions that they're using gpt-oss-120b. There were also some scores shared for gpt-oss-20b, in the 25/50 range. The next "public" model is qwen30b-thinking at 23/50. Competition is limited to 1 H100 (80GB) and 5h runtime for 50 problems.…
The token use chart in the OP release page demonstrates the Qwen issue well.
Token churn does help smaller models on math tasks, but for general purpose stuff it seems to hurt.
Re: Mistral 3 family of models released
#125I use large language models in http://phrasing.app to format data I can retrieve in a consistent skimmable manner. I switched to mistral-3-medium-0525 a few months back after struggling to get gpt-5 to stop producing gibberish. It's been insanely fast, cheap, reliable, and follows formatting instructions to the letter. I was (and still am) super super impressed. Even if it does not hold up in benchmarks, it still out…
Re: Mistral 3 family of models released
#126Earlier quoted context omitted.
Thanks for sharing your use case of the mistral models, which are indeed top-notch ! I had a look at phrasing.app, and while a nice website, I found the copy of "Hand-crafted. Phrasing was designed & developed by humans, for humans." somewhat of a false virtue given your statements here of advanced lllm usage.
I don't see the contention. I do not use llms in the design, development, copywriting, marketing, blogging, or any other aspect of the crafting of the application. I labor over every word, every button, every line of code, every blog post. I would say it is as hand-crafted as something digital can be.
Re: Mistral 3 family of models released
#127Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
Open weight LLMs aren't supposed to "beat" closed models, and they never will. That isn’t their purpose. Their value is as a structural check on the power of proprietary systems; they guarantee a competitive floor. They’re essential to the ecosystem, but they’re not chasing SOTA.
I feel we're only a year or two away from hitting a plateau with the frontier closed models having diminishing returns vs what's "open"
Re: Mistral 3 family of models released
#128Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
Yep, Gemini is my least favorite and I’m convinced that the hype around it isn’t organic because I don’t see the claimed “superiority”, quite the opposite.
Frankly, I don't actually care about or want "general intelligence" -- I want it to make good code, follow instructions, and find bugs. Gemini wasn't bad at the last bit, but wasn't great at the others.
They're all trying to make general purpose AI, but I just want really smart augmentation / tools.
Re: Mistral 3 family of models released
#129Earlier quoted context omitted.
Open weight LLMs aren't supposed to "beat" closed models, and they never will. That isn’t their purpose. Their value is as a structural check on the power of proprietary systems; they guarantee a competitive floor. They’re essential to the ecosystem, but they’re not chasing SOTA.
> Open weight LLMs aren't supposed to "beat" closed models, and they never will. That isn’t their purpose. Do things ever work that way? What if Google did Open source Gemini. Would you say the same? You never know. There's never "supposed" and "purpose" like that.
OpenAI went closed (despite open literally being in the name) once they had the advantage. Meta also is going closed now that they've caught up.
Open-source makes sense to accelerate to catch up, but once ahead, closed will come back to retain advantage.
Re: Mistral 3 family of models released
#130Sad to see they've apparently fully given up on releasing their models via torrent magnet URLs shared on Twitter; those will stay around long after Hugging Face is dead.