Anyone else find that despite Gemini performing best on benches, it's actually still far worse than ChatGPT and Claude? It seems to hallucinate nonsense far more frequently than any of the others. Feels like Google just bench maxes all day every day. As for Mistral, hopefully OSS can eat all of their lunch soon enough.
Mistral 3 family of models released
211–220 of 243 posts
Re: Mistral 3 family of models released
#212I was subscribing to these guys purely to support the EU tech scene. So I was on Pro for about 2 years while using ChatGPT and Claude. Went to actually use it, got a message saying that I missed a payment 8 months previously and thus wasn't allowed to use Pro despite having paid for Pro for the previous 8 months. The lady I contacted in support simply told me to pay the outstanding balance. You would think if you mis…
I'm not sure I understand you correctly, but it seems you had a subscription missed one payment some time ago, but now expect that your subscription works because the missed month was in the past and "you paid for this month"? This sounds like the you expect your subscription to work as an on-demand service? It seems quite obvious that to be able to use a service you would need to be up to date on your payments, that…
That's exactly what it is.
>I'm not sure I understand you correctly,
I understand perfectly well, I don't agree with that approach is the issue.
If I paid for 11/12 months I should get 11/12 months subscription not 1/12 months. They happily just took a years subscription and provided nothing in return. Even if I fixed the outstanding balance they would have provided 2/12 months of service at a cost of 12/12 months of payment.
Re: Mistral 3 family of models released
#213Earlier quoted context omitted.
Thats not the point. Deepmind is not an UK company, its google aka US. Mistral is a real EU based company.
Using US VC dollars. Where their desks are isn’t really important.
And an EU company can't be forced by the US Gov to hand over data.
Re: Mistral 3 family of models released
#214Earlier quoted context omitted.
Yes, of course.
Wow. If all the trillions only produces that small of a diff... that's shocking. That's the sort of knowledge that could pop the bubble.
You can litteraly "improve" your model on LMArena by just adding a bunch of emojis.
Re: Mistral 3 family of models released
#215Earlier quoted context omitted.
This is my experience as well. Mistral models may not be the best according to benchmarks and I don't use them for personal chats or coding, but for simple tasks with pre-defined scope (such as categorization, summarization, etc.) they are the option I choose. I use mistral-small with batch API and it's probably the best cost-efficient option out there.
Did you compare it to gemini-2.0-flash-lite?
Artificial Analysis ranks them close in terms of price (both 0.3 USD/1M tokens) and intelligence (27 / 29 for gemini/mistral), but ranks gemini-2.0-flash-lite higher in terms of speed (189 tokens/s vs. 130).
So they should be interchangeable. Looking forward to testing this.
[0] https://artificialanalysis.ai/?models=o3%2Cgemini-2-5-pro%2C...
Re: Mistral 3 family of models released
#216Earlier quoted context omitted.
Open weight LLMs aren't supposed to "beat" closed models, and they never will. That isn’t their purpose. Their value is as a structural check on the power of proprietary systems; they guarantee a competitive floor. They’re essential to the ecosystem, but they’re not chasing SOTA.
I can attest to Mistral beating OpenAI in my use cases pretty definitively :)
Granted my uses have been programming related. Mistral prints the answer almost immediately and is also completely and utterly hallucinating everything and producing just something that looks like code but could never even compile...
Re: Mistral 3 family of models released
#217I use large language models in http://phrasing.app to format data I can retrieve in a consistent skimmable manner. I switched to mistral-3-medium-0525 a few months back after struggling to get gpt-5 to stop producing gibberish. It's been insanely fast, cheap, reliable, and follows formatting instructions to the letter. I was (and still am) super super impressed. Even if it does not hold up in benchmarks, it still out…
Re: Mistral 3 family of models released
#218Earlier quoted context omitted.
It makes me wonder about the gaps in evaluating LLMs by benchmarks. There almost certainly is overfitting happening which could degrade other use cases. "In practice" evaluation is what inspired the Chatbot Arena right? But then people realized that Chatbot arena over-prioritizes formatting, and maybe sycophancy(?). Makes you wonder what the best evaluation would be. We probably need lots more task-specific models. T…
The best benchmark is one that you build for your use-case. I finally did that for a project and I was not expecting the results. Frontier models are generally "good enough" for most use-cases but if you have something specific you're optimizing for there's probably a more obscure model that just does a better job.
Re: Mistral 3 family of models released
#219Earlier quoted context omitted.
I think people from the US often aren't aware how many companies from the EU simply won't risk losing their data to the providers you have in mind, OpenAI, Anthropic and Google. They simply are no option at all. The company I work for for example, a mid-sized tech business, currently investigates their local hosting options for LLMs. So Mistral certainly will be an option, among the Qwen familiy and Deepseek. Mistral…
Mistral is founded by multiple Meta engineers, no? Funded mostly by US VCs? Hosted primarily on Azure? Do you really have to go out of your way to start calling their competition "data leeches" for out-executing them?
And personally I don't care at all about the performance delta - we are talking about a difference of 6 to at most 12 months here, between closed source SOTA and open weight models.
Re: Mistral 3 family of models released
#220Earlier quoted context omitted.
Some time ago I canceled all my paid subscriptions to chatbots because they are interchangeable so I just rotate between Grok, ChatGPT, Gemini, Deepseek and Mistral. On the API side of things my experience is that the model behaving as expected is the greatest feature. There I also switched to Openrouter instead of paying directly so I can use whatever model fits best. The recent buzz about ad-based chatbot services…
> I guess they hope I forget to cancel. Business model of most subscription based services.