Live data from Hacker News

Notes from the Mistral AI Now Summit

koenvangilst.nl

41–50 of 230 posts

Re: Notes from the Mistral AI Now Summit

#42

OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…

> task focused small models This is tangential: and forgive my ignorance here, but is there an inherent reason why there aren't smaller, focused models from the frontier model providers? I'm thinking something like a software-specific subset of Opus that is the default for use in Claude Code. Smaller, cheaper to deploy and consume, maybe faster.

OpenAI used to make Codex-specific models, but they stopped. What I've gathered from interviews and similar is that training two models isn't worth the (small) lift from having a coding-specific model. You're pre-training on everything anyway, and coding RL is reasonably useful for general-purpose models too.

Re: Notes from the Mistral AI Now Summit

#43
post #3

> BNP Paribas runs Mistral models on-prem for KYC in Belgium, with sensitive data staying within the bank's walls. Abanca is using agent orchestration to handle sensitive customer information at a huge scale (2 million customers in their app). For European companies in regulated industries, this is a good alternative to relying on US hyperscalers. Mistral leaning into on-prem and European-hosted models is very smart.

Yeah but why use mistral on premises instead of Qwen?

[flagged]

Re: Notes from the Mistral AI Now Summit

#44
post #27

I really want Europe to be part of the AI development and research. And I strongly cheered for Mistral. But they are accumulating too much technological delay. This needs to be fixed, otherwise it will turn into yet another proof we are not able to run large tech with good results. Basically any Chinese lab is doing much better. It's not Mistral that created I don't want to say DeepSeek, but MiMo 2.5, Minimax 2.7, an…

https://en.wikipedia.org/wiki/Artificial_Intelligence_Act#Pe... Europe shot itself in the dick with this hastily implemented at the height of mass hysteria bullshit and now no sane company will build anything there. an AI startup in the US or China can be a boy and his computer. in Europe, the boy needs a dozen lawyers. Mistral's sinking into irrelevancy despite the head start they had, the very promising early model…

So you're saying AI models should be allowed to freely "manipulate human behavior"?

Re: Notes from the Mistral AI Now Summit

#45
post #26
post #16

Earlier quoted context omitted.

Nobody trying to compete with Google, OpenAI, and Anthropic should be playing the small models / local models game. Foundation model labs should be building very large reasoning models, then leaving it to the community to distill them down. You can't scale a small model up, but you can scale a small model down. I'm convinced the only way we'll have a seat at the table in the future and avoid total runaway takeoff is…

I thought distillation meant small models don't have to compete with the big models and can always eventually achieve close parity, but it's just a matter of time to do the distillation? (i.e. how much lag do you want to live with) Am I oversimplifying?

There is likely a theoretical limit to how much intelligence you can pack into a model of a given size (especially when stretching that over a large input context size).

Our evals are pretty complex so we only recently started testing ~30B class models, which are now becoming quite smart (on par with the frontier from 1 year ago). Mistral is far behind, but I'm rooting for them.

Data at https://gertlabs.com/rankings

Re: Notes from the Mistral AI Now Summit

#46
post #3

> BNP Paribas runs Mistral models on-prem for KYC in Belgium, with sensitive data staying within the bank's walls. Abanca is using agent orchestration to handle sensitive customer information at a huge scale (2 million customers in their app). For European companies in regulated industries, this is a good alternative to relying on US hyperscalers. Mistral leaning into on-prem and European-hosted models is very smart.

Yeah but why use mistral on premises instead of Qwen?

Please don't run Chinese models for KYC operations.

Re: Notes from the Mistral AI Now Summit

#47
post #44

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Artificial_Intelligence_Act#Pe... Europe shot itself in the dick with this hastily implemented at the height of mass hysteria bullshit and now no sane company will build anything there. an AI startup in the US or China can be a boy and his computer. in Europe, the boy needs a dozen lawyers. Mistral's sinking into irrelevancy despite the head start they had, the very promising early model…

So you're saying AI models should be allowed to freely "manipulate human behavior"?

The problem is that statement is a bit too open to interpretation. Ever had Claude piss you off by being stupid and talking in circles? Sounds like manipulation of human behavior!

Re: Notes from the Mistral AI Now Summit

#48
post #40

OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…

I agree. I am a paying Le Chat Pro user, really rooting for a European alternative. But the quality difference between Mistral and the frontier labs is growing too big to ignore. It’s worrying to me that they didn’t talk much about new models at the conference, because that is really where their focus should be IMHO. I am wondering what is keeping them back, though: Money? Compute? Skills? Training data? My fear is t…

My theory with no insider information: it’s a little of all of the above, but mostly money. To some extent, you can dig yourself out of a data hole with RL and a lot of compute. And you can buy a lot of compute and some data with a lot of money. Big labs have been operating in this regime for a while and it’s one of the drivers behind their costs beyond just scaling the weights and doing the actual training. Mistral just doesn’t have access to this level of compute or the money to try and muscle their way in.

Re: Notes from the Mistral AI Now Summit

#49

OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…

I don't agree that they are falling behind. Using both chat and cli I get what I need and it's comparable to "sota" when I compare.

Re: Notes from the Mistral AI Now Summit

#50

Earlier quoted context omitted.

> task focused small models This is tangential: and forgive my ignorance here, but is there an inherent reason why there aren't smaller, focused models from the frontier model providers? I'm thinking something like a software-specific subset of Opus that is the default for use in Claude Code. Smaller, cheaper to deploy and consume, maybe faster.

OpenAI used to make Codex-specific models, but they stopped. What I've gathered from interviews and similar is that training two models isn't worth the (small) lift from having a coding-specific model. You're pre-training on everything anyway, and coding RL is reasonably useful for general-purpose models too.

Interesting. I'd have guessed there would be meaningful opex benefits to serving smaller models.
Post reply on HN