Live data from Hacker News

Notes from the Mistral AI Now Summit

koenvangilst.nl

91–100 of 230 posts

Re: Notes from the Mistral AI Now Summit

#91
post #73

Earlier quoted context omitted.

They can get what, 1B euros? 10B when everyone loses their mind? This doesn’t buy nearly enough compute nowadays. Meanwhile, Anthropic and OpenAI have investors practically begging them to let them buy this much equity at mind-bogging valuations.

[flagged]

You think they're intentionally being bad because they can't manage to pump $65B into a startup on a whim...?

Re: Notes from the Mistral AI Now Summit

#92

OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…

> a decent proxy would be to build models that get the r/localLlama crowd excited

I don’t really disagree with your post, but this is not exactly right. That subreddit seems to go from hype train to hype train every week, I haven’t found anything really insightful in it for quite a while now.

Re: Notes from the Mistral AI Now Summit

#94

Earlier quoted context omitted.

OpenAI used to make Codex-specific models, but they stopped. What I've gathered from interviews and similar is that training two models isn't worth the (small) lift from having a coding-specific model. You're pre-training on everything anyway, and coding RL is reasonably useful for general-purpose models too.

Interesting. I'd have guessed there would be meaningful opex benefits to serving smaller models.

[deleted]

Re: Notes from the Mistral AI Now Summit

#95
post #40

OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…

I agree. I am a paying Le Chat Pro user, really rooting for a European alternative. But the quality difference between Mistral and the frontier labs is growing too big to ignore. It’s worrying to me that they didn’t talk much about new models at the conference, because that is really where their focus should be IMHO. I am wondering what is keeping them back, though: Money? Compute? Skills? Training data? My fear is t…

> I am wondering what is keeping them back, though: Money? Compute? Skills? Training data?

Not ruthless enough and no backing by a corrupt govt administration that has no morals but focuses on self-enrichment instead.

Might sound drastic but I think that's actually closer to the truth thn everbody likes to admit.

> My fear is that you are really only getting really good models by training on very dubious data (outputs from the frontier models etc) and that Mistral is too European and too enterprisey to take those risks.

Exactly.

Re: Notes from the Mistral AI Now Summit

#96

Earlier quoted context omitted.

Because they distill

it doesn't matter the reason. This is a race and nobody will care or remember how the winners got there. Mistral looks like it's fading away to irrelevance unless they can play alongside the similar sized models, or have some unique advantage other than being in Europe, for Europe. I was really excited for them back when they were startup that had the biggest European venture round ever. This space will have a few wi…

It’s not that I don’t agree with you, I am just pointing out why it’s hard to catch up to scaling laws given the European economic (capital) and political (US would be upset if they found out Europeans distill) constraints. China is only bound by economic constraints.

Re: Notes from the Mistral AI Now Summit

#97
post #87

Earlier quoted context omitted.

Don’t they supposedly have a huge amount of EU support? Or at least there’s been a lot of noise about that.

It's a bit strange, but a huge handout from the EU/France and a huge AI lab investment round are different orders of magnitude. The necessary sums are just not politically possible. How do you sell spending the equivalent of ten USS Gerald Fords on a start-up? You don't.

And a lot of the "funding" is through mutual deals with MSFT, Nvidia, etc. The Europeans have none of that and would need to pay in actual cash.

Re: Notes from the Mistral AI Now Summit

#98
post #59

OK, I'm 100% rooting for both Mistral and task focused small models. But Mistral has fall really far behind since 2025Q3. It seems they can't get good reasoning models working at even medium context sizes, which is necessary to be at the table right now. Gemma4 and Qwen3.6 are currently best in the small size; Mistral's "small" model has ~4x the parameter count at 120B and isn't even competing with models a quarter i…

Yeah. I run LLM models locally and for me 22B-32B is the largest I'm willing to invest in trying out. Even though Mistral 4 has 6B active parameters per token (allowing 3-3.5 per token parameters to be loaded on a 4090), the ~240GB download + storage is pushing the limits of being able to try this out locally, especially if you are downloading and evaluating multiple models. It also makes it harder for other people t…

I think machines like the DGX Spark are about to become a lot more common/popular. It’s big enough to run sparse 150-250B MoEs with enough throughout for a single user. Deepseek v4 Flash is #1 (in terms of usage) on OpenRouter because it’s good enough to be useful. You can run it on a Spark (though it runs better across 2, which is getting up there in cost)

Re: Notes from the Mistral AI Now Summit

#100
post #56

Earlier quoted context omitted.

DeepSeek is both cheaper and better than Mistral.

Because they distill

I feel like there's an implication here that distillation is a problem but I don't understand what you mean. I thought distillation was generating text from a model and then training another model on it. Is the something unethical in that? You're paying the API costs to generate the tokens, right?

Or I guess more to the point: is this something frontier labs have said is (or tried to paint at any rate) problematic? This feels like an "out of the loop" situation because I've only ever heard "distillation" with a positive connotation before.

Post reply on HN