Earlier quoted context omitted.
Some time ago I canceled all my paid subscriptions to chatbots because they are interchangeable so I just rotate between Grok, ChatGPT, Gemini, Deepseek and Mistral. On the API side of things my experience is that the model behaving as expected is the greatest feature. There I also switched to Openrouter instead of paying directly so I can use whatever model fits best. The recent buzz about ad-based chatbot services…
Maybe give Perplexity a shot? It has Grok, ChatGPT, Gemini, Kimi K2, I dont think it has Mistral unfortunately.
Mistral 3 family of models released
151–160 of 243 posts
Re: Mistral 3 family of models released
#152Re: Mistral 3 family of models released
#153Earlier quoted context omitted.
1. Big problem 2. ASML was propped up by ASM and Philips, stepping in as "VCs"
For VC don't you need a lot of capital and people with too much money? Isn't that then a chicken and egg?
No. VC’s historical capital has come from institutional investors. Pensions. Endowments. Foundations.
Re: Mistral 3 family of models released
#154Earlier quoted context omitted.
Thats not the point. Deepmind is not an UK company, its google aka US. Mistral is a real EU based company.
Using US VC dollars. Where their desks are isn’t really important.
Re: Mistral 3 family of models released
#155Earlier quoted context omitted.
The lack of the comparison (which absolutely was done), tells you exactly what you need to know.
I think people from the US often aren't aware how many companies from the EU simply won't risk losing their data to the providers you have in mind, OpenAI, Anthropic and Google. They simply are no option at all. The company I work for for example, a mid-sized tech business, currently investigates their local hosting options for LLMs. So Mistral certainly will be an option, among the Qwen familiy and Deepseek. Mistral…
Re: Mistral 3 family of models released
#156Re: Mistral 3 family of models released
#157Re: Mistral 3 family of models released
#158Sad to see they've apparently fully given up on releasing their models via torrent magnet URLs shared on Twitter; those will stay around long after Hugging Face is dead.
How does HF manage to serve such big files?
Re: Mistral 3 family of models released
#159Earlier quoted context omitted.
I have a need to remove loose "signature" lines from the last 10% of a tremendous e-mail dataset. Based on your experience, how do you think mistral-3-medium-0525 would do?
What's your acceptable error rate? Honestly ministral would probably be sufficient if you can tolerate a small failure rate. I feel like medium would be overkill. But I'm no expert. I can't say I've used mistral much outside of my own domain.
Re: Mistral 3 family of models released
#160I use large language models in http://phrasing.app to format data I can retrieve in a consistent skimmable manner. I switched to mistral-3-medium-0525 a few months back after struggling to get gpt-5 to stop producing gibberish. It's been insanely fast, cheap, reliable, and follows formatting instructions to the letter. I was (and still am) super super impressed. Even if it does not hold up in benchmarks, it still out…
It makes me wonder about the gaps in evaluating LLMs by benchmarks. There almost certainly is overfitting happening which could degrade other use cases. "In practice" evaluation is what inspired the Chatbot Arena right? But then people realized that Chatbot arena over-prioritizes formatting, and maybe sycophancy(?). Makes you wonder what the best evaluation would be. We probably need lots more task-specific models. T…
Also, we should be aware of people cynically playing into that bias to try to advertise their app, like OP who has managed to spam a link in the first line of a top comment on this popular front page article by telling the audience exactly what they want to hear ;)