Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
Mistral 3 family of models released
111–120 of 243 posts
Re: Mistral 3 family of models released
#112Re: Mistral 3 family of models released
#113Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
Re: Mistral 3 family of models released
#114Re: Mistral 3 family of models released
#115I use large language models in http://phrasing.app to format data I can retrieve in a consistent skimmable manner. I switched to mistral-3-medium-0525 a few months back after struggling to get gpt-5 to stop producing gibberish. It's been insanely fast, cheap, reliable, and follows formatting instructions to the letter. I was (and still am) super super impressed. Even if it does not hold up in benchmarks, it still out…
Thanks for sharing your use case of the mistral models, which are indeed top-notch ! I had a look at phrasing.app, and while a nice website, I found the copy of "Hand-crafted. Phrasing was designed & developed by humans, for humans." somewhat of a false virtue given your statements here of advanced lllm usage.
I labor over every word, every button, every line of code, every blog post. I would say it is as hand-crafted as something digital can be.
Re: Mistral 3 family of models released
#116Earlier quoted context omitted.
The lack of the comparison (which absolutely was done), tells you exactly what you need to know.
I think people from the US often aren't aware how many companies from the EU simply won't risk losing their data to the providers you have in mind, OpenAI, Anthropic and Google. They simply are no option at all. The company I work for for example, a mid-sized tech business, currently investigates their local hosting options for LLMs. So Mistral certainly will be an option, among the Qwen familiy and Deepseek. Mistral…
Funded mostly by US VCs?
Hosted primarily on Azure?
Do you really have to go out of your way to start calling their competition "data leeches" for out-executing them?
Re: Mistral 3 family of models released
#117Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
It's also slower than both Opus 4.5 and Sonnet.
Re: Mistral 3 family of models released
#118Earlier quoted context omitted.
Thats not the point. Deepmind is not an UK company, its google aka US. Mistral is a real EU based company.
Using US VC dollars. Where their desks are isn’t really important.
The cloud act and the current US administration doing things like sanctioning the ICC demonstrate why the locations of those desks is important.
Re: Mistral 3 family of models released
#119Re: Mistral 3 family of models released
#120Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
In prior posts you oddly attack "Palantir-partnered Anthropic" as well.
Are things that grim at OpenAI that this sort of FUD is necessary? I mean, I know they're doing the whole code red thing, but I guarantee that posting nonsense like this on HN isn't the way.