Live data from Hacker News

Mistral Medium 3.5

mistral.ai

171–180 of 248 posts

Re: Mistral Medium 3.5

#171

This is a very interesting strategy that might pay off. This model is a very good option for enterprise self host. I would argue a lot of companies are VRAM constrained rather than compute constrained. You could fit 4-5 running instances on one H100 cluster where you can only fit 1-2 Kimi K2 or GLM5.

This is 128B dense though. the K/V cache on long context is going to be massive

With turbo quant, you would reduce it by over 6X.

Re: Mistral Medium 3.5

#172
post #166

Earlier quoted context omitted.

And GPT-2 1.5B was considered too dangerous to release. They were perhaps right.

considered that by OpenAI for marketing purposes that is But yes, perhaps it would have been better for all of us if they haven't.

In lockstep over the past month, a subset of people, un-labelable, unprompted, share this train of thought:

- Mythos wasn't released widely.

- But Anthropic shared info on it and said it was dangerous.

- Anthropic is a company.

- Companies like money.

- Therefore Mythos is marketing hype.

- Remember GPT-2? That also wasn't released. They said it was dangerous.

- But, GPT-3, GPT-4, GPT-5, etc. were released.

- Therefore GPT-2 being dangerous was marketing hype.

I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared.

It's Not Even Wrong, in the Pauli sense: they weren't selling anything! They weren't raising funding! What were they marketing!?

And there's a lot more elided from history, ex. they didn't have an API yet.

GPT-3 was released, a year or two later, and did have an API. But, no one used it, it wasn't good enough yet. And they did treat it as dangerous, it was wildly over-the-top manually monitored for anything resembling not-intended-use. I got permanently suspended for using the word "twink"

Re: Mistral Medium 3.5

#174
post #78
post #51

Earlier quoted context omitted.

DeepMind, which is headquartered in London, probably had a significant role in the development of the Gemini and Gemma models. Yes, it might be a problem that the UK allows companies like this to be bought up by foreign countries.

Without Google’s funding its not obvious i DeepMind would have went anywhere. Unless the moved to US for funding while keeping a back office in the UK. It’s strange to expect anything significant to come out from Europe when VCs there are either very risk averse and/or don’t have enough cash to begin with. It’s not like government or EU funding can replace that since its almost always wasted or missdirected

It’s a company containing such remarkable talent that I’m sure they would not have run into significant issues raising capital on international markets.

It’s not like VCs are only allowed to invest in companies in their own country.

Re: Mistral Medium 3.5

#175
post #166

Earlier quoted context omitted.

considered that by OpenAI for marketing purposes that is But yes, perhaps it would have been better for all of us if they haven't.

In lockstep over the past month, a subset of people, un-labelable, unprompted, share this train of thought: - Mythos wasn't released widely. - But Anthropic shared info on it and said it was dangerous. - Anthropic is a company. - Companies like money. - Therefore Mythos is marketing hype. - Remember GPT-2? That also wasn't released. They said it was dangerous. - But, GPT-3, GPT-4, GPT-5, etc. were released. - Therefo…

> I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared.

That's not what I am saying.

It's not that GPT-2 not being released was marketing hype, it's that OpenAI themselves claiming it's too dangerous to release specifically, implying it's close to AGI, (or something like that), was marketing hype.

Re: Mistral Medium 3.5

#176
post #163
post #13

Earlier quoted context omitted.

This is the bar for Europe, huh?

I mean, at least we're not melting the planet trying to predict the next token that sounds about right.

Europeans use AI as much as anyone else.

Re: Mistral Medium 3.5

#177
post #175

Earlier quoted context omitted.

In lockstep over the past month, a subset of people, un-labelable, unprompted, share this train of thought: - Mythos wasn't released widely. - But Anthropic shared info on it and said it was dangerous. - Anthropic is a company. - Companies like money. - Therefore Mythos is marketing hype. - Remember GPT-2? That also wasn't released. They said it was dangerous. - But, GPT-3, GPT-4, GPT-5, etc. were released. - Therefo…

> I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared. That's not what I am saying. It's not that GPT-2 not being released was marketing hype, it's that OpenAI themselves claiming it's too dangerous to release specifically, implying it's close to AGI, (or something like that), was marketing hype.

That may sound more defensible to you, but its even more detached from reality. I feel very old right now because I actually read the thing at the time, but setting that aside, do you really think anyone thought or said GPT-2 was AGI?

I don't think you do.

I only mention reading it because that would clear it up, and you seem interested, and your parenthetical indicates A) you're aware you're claiming something a bit silly and B) you don't know what was actually said.

Re: Mistral Medium 3.5

#178
post #13

Earlier quoted context omitted.

This is the bar for Europe, huh?

[flagged]

The fact that this comment is still up hours later but my comment below participating in the discussion got flagged should tell one everything they need to know about the intellectual rigor here.

Re: Mistral Medium 3.5

#179
post #176
post #163

Earlier quoted context omitted.

I mean, at least we're not melting the planet trying to predict the next token that sounds about right.

Europeans use AI as much as anyone else.

Yes, but it would seem that Chinese models are much more efficiently trained than the US ones, (i.e. with fewer resources).

Europe doesn't invest nowhere near as much as the US does into tech, so we need to either figure out how to be at least as, and hopefully more, efficient as the Chinese models are (at least in terms of training) or there's little point in trying.

I suspect this is one of the reasons why Mistral's models are somewhat struggling; i.e. US style training costs, but nowhere near as much cash as OpenAI/Anthopic have.

There are multiple European Google alternatives as well for example, but being 80% as good just doesn't cut it. Chinese models win because they are 95-98% as good as the SotA US ones but at a fraction of the cost.

Re: Mistral Medium 3.5

#180
post #52

I'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size. GLM 5.1 is an excellent model, but even at Q4 you're looking at ~400GB. Kimi K2.5 is really good too, and at Q4 quantization you're looking at almost ~600GB. This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM…

“This beats the latest Sonnet while running locally”

Not really.

- The benchmarks are based on F8_E4M3 and you’re not running that on any Mac.

- Sonnet has a 1M token context window. This is 256k but again you’re probably not even getting that locally.

- Sonnet is fast over the wire. This is going to be much slower.

Post reply on HN