Regardless of the business. Their website design is :chefs-kiss https://mistral.ai/
Notes from the Mistral AI Now Summit
41–50 of 230 posts
Re: Notes from the Mistral AI Now Summit
#42OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…
> task focused small models This is tangential: and forgive my ignorance here, but is there an inherent reason why there aren't smaller, focused models from the frontier model providers? I'm thinking something like a software-specific subset of Opus that is the default for use in Claude Code. Smaller, cheaper to deploy and consume, maybe faster.
Re: Notes from the Mistral AI Now Summit
#43> BNP Paribas runs Mistral models on-prem for KYC in Belgium, with sensitive data staying within the bank's walls. Abanca is using agent orchestration to handle sensitive customer information at a huge scale (2 million customers in their app). For European companies in regulated industries, this is a good alternative to relying on US hyperscalers. Mistral leaning into on-prem and European-hosted models is very smart.
Yeah but why use mistral on premises instead of Qwen?
Re: Notes from the Mistral AI Now Summit
#44I really want Europe to be part of the AI development and research. And I strongly cheered for Mistral. But they are accumulating too much technological delay. This needs to be fixed, otherwise it will turn into yet another proof we are not able to run large tech with good results. Basically any Chinese lab is doing much better. It's not Mistral that created I don't want to say DeepSeek, but MiMo 2.5, Minimax 2.7, an…
https://en.wikipedia.org/wiki/Artificial_Intelligence_Act#Pe... Europe shot itself in the dick with this hastily implemented at the height of mass hysteria bullshit and now no sane company will build anything there. an AI startup in the US or China can be a boy and his computer. in Europe, the boy needs a dozen lawyers. Mistral's sinking into irrelevancy despite the head start they had, the very promising early model…
Re: Notes from the Mistral AI Now Summit
#45Earlier quoted context omitted.
Nobody trying to compete with Google, OpenAI, and Anthropic should be playing the small models / local models game. Foundation model labs should be building very large reasoning models, then leaving it to the community to distill them down. You can't scale a small model up, but you can scale a small model down. I'm convinced the only way we'll have a seat at the table in the future and avoid total runaway takeoff is…
I thought distillation meant small models don't have to compete with the big models and can always eventually achieve close parity, but it's just a matter of time to do the distillation? (i.e. how much lag do you want to live with) Am I oversimplifying?
Our evals are pretty complex so we only recently started testing ~30B class models, which are now becoming quite smart (on par with the frontier from 1 year ago). Mistral is far behind, but I'm rooting for them.
Data at https://gertlabs.com/rankings
Re: Notes from the Mistral AI Now Summit
#46> BNP Paribas runs Mistral models on-prem for KYC in Belgium, with sensitive data staying within the bank's walls. Abanca is using agent orchestration to handle sensitive customer information at a huge scale (2 million customers in their app). For European companies in regulated industries, this is a good alternative to relying on US hyperscalers. Mistral leaning into on-prem and European-hosted models is very smart.
Yeah but why use mistral on premises instead of Qwen?
Re: Notes from the Mistral AI Now Summit
#47Earlier quoted context omitted.
https://en.wikipedia.org/wiki/Artificial_Intelligence_Act#Pe... Europe shot itself in the dick with this hastily implemented at the height of mass hysteria bullshit and now no sane company will build anything there. an AI startup in the US or China can be a boy and his computer. in Europe, the boy needs a dozen lawyers. Mistral's sinking into irrelevancy despite the head start they had, the very promising early model…
So you're saying AI models should be allowed to freely "manipulate human behavior"?
Re: Notes from the Mistral AI Now Summit
#48OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…
I agree. I am a paying Le Chat Pro user, really rooting for a European alternative. But the quality difference between Mistral and the frontier labs is growing too big to ignore. It’s worrying to me that they didn’t talk much about new models at the conference, because that is really where their focus should be IMHO. I am wondering what is keeping them back, though: Money? Compute? Skills? Training data? My fear is t…
Re: Notes from the Mistral AI Now Summit
#49OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…
Re: Notes from the Mistral AI Now Summit
#50Earlier quoted context omitted.
> task focused small models This is tangential: and forgive my ignorance here, but is there an inherent reason why there aren't smaller, focused models from the frontier model providers? I'm thinking something like a software-specific subset of Opus that is the default for use in Claude Code. Smaller, cheaper to deploy and consume, maybe faster.
OpenAI used to make Codex-specific models, but they stopped. What I've gathered from interviews and similar is that training two models isn't worth the (small) lift from having a coding-specific model. You're pre-training on everything anyway, and coding RL is reasonably useful for general-purpose models too.