Live data from Hacker News

Mistral 3 family of models released

mistral.ai

81–90 of 243 posts

Re: Mistral 3 family of models released

#81
post #65

Earlier quoted context omitted.

Thats not the point. Deepmind is not an UK company, its google aka US. Mistral is a real EU based company.

Using US VC dollars. Where their desks are isn’t really important.

Currency is interchangeable. Location might not be.

Re: Mistral 3 family of models released

#82

Earlier quoted context omitted.

1491 vs 1418 ELO means the stronger model wins about 60% of the time.

Probably naive questions: Does that also mean that Gemini-3 (the top ranked model) loses to mistral 3 40% of the time? Does that make Gemini 1.5x better, or mistral 2/3rd as good as Gemini, or can we not quantify the difference like that?

Yes, of course.

Re: Mistral 3 family of models released

#84
post #47
post #40

Europe's bright star has been quiet for a while, great to see them back and good to see them come back to Open Source light with Apache 2.0 licenses - they're too far from the SOTA pack that exclusive/proprietary models would work in their favor. Mistral had the best small models on consumer GPUs for a while, hopefully Ministral 14B lives up to their benchmarks.

All thanks to the US VCs that acutally have money to fund Mistral's entire business. Had they gone to the EU, Mistral would have gotten a miniscule grant from the EU to train their AI models.

Mistral biggest investor is asml, although it became so later than other vcs

Re: Mistral 3 family of models released

#85
Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.

Re: Mistral 3 family of models released

#86
post #73

Earlier quoted context omitted.

Yes. I spent about 3 days trying to optimize the prompt to get gpt-5 to not produce gibberish, to no avail. Completions took several minutes, had an above 50% timeout rate (with a 6 minute timeout mind you), and after retrying they still would return gibberish about 15% of the time (12% on one task, 20% on another task). I then tried multiple models, and they all failed in spectacular ways. Only Grok and Mistral had…

Hard to gauge what gibberish is without an example of the data and what you prompted the LLM with.

If you wanted examples, you needed only ask :)

These are screenshots from that week: https://x.com/barrelltech/status/1995900100174880806

I'm not going to share the prompt because (1) it's very long (2) there were dozens of variations and (3) it seems like poor business practices to share the most indefensible part of your business online XD

Re: Mistral 3 family of models released

#87

Earlier quoted context omitted.

The lack of the comparison (which absolutely was done), tells you exactly what you need to know.

They're comparing against open weights models that are roughly a month away from the frontier. Likely there's an implicit open-weights political stance here. There are also plenty of reasons not to use proprietary US models for comparison: The major US models haven't been living up to their benchmarks; their releases rarely include training & architectural details; they're not terribly cost effective; they often fail…

Scale AI wrote a paper a year ago comparing various models performance on benchmarks to performance on similar but held-out questions. Generally the closed source models performed better, and Mistral came out looking pretty badly: https://arxiv.org/pdf/2405.00332

Re: Mistral 3 family of models released

#88
post #4

I still don't understand what the incentive is for releasing genuinely good model weights. What makes sense however is OpenAI releasing a somewhat generic model like gpt-oss that games the benchmarks just for PR. Or some Chinese companies doing the same to cut the ground from under the feet of American big tech. Are we really hopeful we'll still get decent open weights models in the future?

Google games benchmarks more than anyone, hence Gemini's strong bench lead. In reality though, it's still garbage for general usage.

Re: Mistral 3 family of models released

#90
post #85

Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.

No, I've been using Gemini for help while learning / building my onprem k8s cluster and it has been almost spotless.

Granted, this is a subject that is very well present in the training data but still.

Post reply on HN