This is a very interesting strategy that might pay off. This model is a very good option for enterprise self host. I would argue a lot of companies are VRAM constrained rather than compute constrained. You could fit 4-5 running instances on one H100 cluster where you can only fit 1-2 Kimi K2 or GLM5.
This is 128B dense though. the K/V cache on long context is going to be massive
Mistral Medium 3.5
171–180 of 248 posts
Re: Mistral Medium 3.5
#172Earlier quoted context omitted.
And GPT-2 1.5B was considered too dangerous to release. They were perhaps right.
considered that by OpenAI for marketing purposes that is But yes, perhaps it would have been better for all of us if they haven't.
- Mythos wasn't released widely.
- But Anthropic shared info on it and said it was dangerous.
- Anthropic is a company.
- Companies like money.
- Therefore Mythos is marketing hype.
- Remember GPT-2? That also wasn't released. They said it was dangerous.
- But, GPT-3, GPT-4, GPT-5, etc. were released.
- Therefore GPT-2 being dangerous was marketing hype.
I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared.
It's Not Even Wrong, in the Pauli sense: they weren't selling anything! They weren't raising funding! What were they marketing!?
And there's a lot more elided from history, ex. they didn't have an API yet.
GPT-3 was released, a year or two later, and did have an API. But, no one used it, it wasn't good enough yet. And they did treat it as dangerous, it was wildly over-the-top manually monitored for anything resembling not-intended-use. I got permanently suspended for using the word "twink"
Re: Mistral Medium 3.5
#173Re: Mistral Medium 3.5
#174Earlier quoted context omitted.
DeepMind, which is headquartered in London, probably had a significant role in the development of the Gemini and Gemma models. Yes, it might be a problem that the UK allows companies like this to be bought up by foreign countries.
Without Google’s funding its not obvious i DeepMind would have went anywhere. Unless the moved to US for funding while keeping a back office in the UK. It’s strange to expect anything significant to come out from Europe when VCs there are either very risk averse and/or don’t have enough cash to begin with. It’s not like government or EU funding can replace that since its almost always wasted or missdirected
It’s not like VCs are only allowed to invest in companies in their own country.
Re: Mistral Medium 3.5
#175Earlier quoted context omitted.
considered that by OpenAI for marketing purposes that is But yes, perhaps it would have been better for all of us if they haven't.
In lockstep over the past month, a subset of people, un-labelable, unprompted, share this train of thought: - Mythos wasn't released widely. - But Anthropic shared info on it and said it was dangerous. - Anthropic is a company. - Companies like money. - Therefore Mythos is marketing hype. - Remember GPT-2? That also wasn't released. They said it was dangerous. - But, GPT-3, GPT-4, GPT-5, etc. were released. - Therefo…
That's not what I am saying.
It's not that GPT-2 not being released was marketing hype, it's that OpenAI themselves claiming it's too dangerous to release specifically, implying it's close to AGI, (or something like that), was marketing hype.
Re: Mistral Medium 3.5
#176Re: Mistral Medium 3.5
#177Earlier quoted context omitted.
In lockstep over the past month, a subset of people, un-labelable, unprompted, share this train of thought: - Mythos wasn't released widely. - But Anthropic shared info on it and said it was dangerous. - Anthropic is a company. - Companies like money. - Therefore Mythos is marketing hype. - Remember GPT-2? That also wasn't released. They said it was dangerous. - But, GPT-3, GPT-4, GPT-5, etc. were released. - Therefo…
> I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared. That's not what I am saying. It's not that GPT-2 not being released was marketing hype, it's that OpenAI themselves claiming it's too dangerous to release specifically, implying it's close to AGI, (or something like that), was marketing hype.
I don't think you do.
I only mention reading it because that would clear it up, and you seem interested, and your parenthetical indicates A) you're aware you're claiming something a bit silly and B) you don't know what was actually said.
Re: Mistral Medium 3.5
#178Re: Mistral Medium 3.5
#179Earlier quoted context omitted.
I mean, at least we're not melting the planet trying to predict the next token that sounds about right.
Europeans use AI as much as anyone else.
Europe doesn't invest nowhere near as much as the US does into tech, so we need to either figure out how to be at least as, and hopefully more, efficient as the Chinese models are (at least in terms of training) or there's little point in trying.
I suspect this is one of the reasons why Mistral's models are somewhat struggling; i.e. US style training costs, but nowhere near as much cash as OpenAI/Anthopic have.
There are multiple European Google alternatives as well for example, but being 80% as good just doesn't cut it. Chinese models win because they are 95-98% as good as the SotA US ones but at a fraction of the cost.
Re: Mistral Medium 3.5
#180I'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size. GLM 5.1 is an excellent model, but even at Q4 you're looking at ~400GB. Kimi K2.5 is really good too, and at Q4 quantization you're looking at almost ~600GB. This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM…
Not really.
- The benchmarks are based on F8_E4M3 and you’re not running that on any Mac.
- Sonnet has a 1M token context window. This is 256k but again you’re probably not even getting that locally.
- Sonnet is fast over the wire. This is going to be much slower.