Live data from Hacker News

Magistral — the first reasoning model by Mistral AI

mistral.ai

211–220 of 444 posts

Re: Magistral — the first reasoning model by Mistral AI

#211
post #11

Benchmarks suggest this model loses to Deepseek-R1 in every one-shot comparison. Considering they were likely not even pitting it against the newer R1 version (no mention of that in the article) and at more than double the cost, this looks like the best AI company in the EU is struggling to keep up with the state-of-the-art.

Jm2c but I feel conflicted about this arms race. You can be 6/12 months later, and have not burned tens of billions compared to the best in class, I see it an engineering win. I absolutely understand those that say "yeah, but customers will only use the best", I see it, but is market share of forever money losing businesses that valuable?

A similar sentiment existed for a long time about Uber and now they're very profitable and own their market. It was worth the burn to capture the market. Who says OpenAI can't roll over to profitable at a stable scale? Conquer the market, hike the price to $29.95 (family account, no ads; $19.95 individual account with ads; etc etc). To say nothing of how they can branch out in terms of being the interaction point that replaces the search box. The advertising value of owning the land that OpenAI is taking is well over $100 billion in annual revenue. Amazon's retail business is terrible, their ad business is fantastic. As OpenAI bolts on an ad product their margin potential will skyrocket and the cost side will be modest in comparison.

Over the coming years it won't be possible to stay a mere 6-12 months behind as the costs to build and maintain the AI super-infrastructure keeps climbing. It'll become a guaranteed implosion scenario. Winning will provide the ongoing immense resources needed to keep pushing up the hill forever. Everybody else - except a few - will fall away. The same outcome took place in search. Anybody spot Lycos, Excite, Hotbot, AltaVista around? It costs an enormous amount of money to try to keep up with Google (Bing, Baidu, Yandex) in search and scale it. This will be an even more brutal example of that, as the costs are even higher to scale.

The only way Mistral survives is if they're heavily subsidized directly by European states.

Re: Magistral — the first reasoning model by Mistral AI

#212
post #119

Here are my notes on trying this out locally via Ollama and via their API (and the llm-mistral plugin) too: https://simonwillison.net/2025/Jun/10/magistral/

Hi Simon, What's the huge difference between the two pelicans riding bicycles? Was one running locally the small version vs the pretty good one running the bigger one thru the API? Thanks, Morgan

Yes, the bad one was Mistral Small running locally, the better one was Mistral Medium via their API.

Re: Magistral — the first reasoning model by Mistral AI

#213
post #204

Earlier quoted context omitted.

There have been 11 mass shootings in the US in the last 7 days so I don't think this disgusting competition is one you're likely to win.

Nobody is claiming the US has less mass shootings. It's just pointless whataboutism in a conversation (economic strategy) that has nothing to do with it.

Ah good, I thought you were trying to imply there is an equivalent problem in the EU. Which would seem to be intentionally dense of course.

Re: Magistral — the first reasoning model by Mistral AI

#214
post #139

Earlier quoted context omitted.

I didn't downvote. T the problem with the paper is that it asks the model to output all moves for, say, 15 disks 2 ^ 15 - 1 = 32767 32767 moves in a single prompt. That's not testing reasoning. That’s testing whether the model can emit a huge structured output without error, under a context window limit. The authors then treat failure to reproduce this entire sequence as evidence that the model can't reason. But that…

No worries, I wasn’t saying to you directly. I agree 15 disks is very difficult for a human, probably on a sheer stamina level; but I managed to do 8 in about 15 minutes by playing around (I.e. no practice). They do state that there is a massive drop in performance at this point.

Remember that with Towers of Hanoi every extra disk doubles the number of moves required. So 15 discs is 128x more moves. If you did eight in 15m then fifteen would take you 32 hours.

Re: Magistral — the first reasoning model by Mistral AI

#215
post #11

Benchmarks suggest this model loses to Deepseek-R1 in every one-shot comparison. Considering they were likely not even pitting it against the newer R1 version (no mention of that in the article) and at more than double the cost, this looks like the best AI company in the EU is struggling to keep up with the state-of-the-art.

Jm2c but I feel conflicted about this arms race. You can be 6/12 months later, and have not burned tens of billions compared to the best in class, I see it an engineering win. I absolutely understand those that say "yeah, but customers will only use the best", I see it, but is market share of forever money losing businesses that valuable?

Indeed, and with the technology plateau-ing, being 6-12 months late with less debt is just long term thinking.

Also, Europe being in the race is a big deal for consumers.

Re: Magistral — the first reasoning model by Mistral AI

#216
Their OCR model was really well hyped and coincidentally came out at the time I had a batch of 600 page pdfs to OCR. They were all monospace text just for some reason the OCR was missing.

I tried it, 80% of the "text" was recognised as images and output as whitespace so most of it was empty. It was much much worse than tesseract.

A month later I got the bill for that crap and deleted my account.

Maybe this is better but I'm over hype marketing from mistral

Re: Magistral — the first reasoning model by Mistral AI

#217

Earlier quoted context omitted.

This is not true. Government workers or factory workers can limit to 35h (with some salary loss or days off loss), but else than that (especially in tech) it is very competitive and working 50 hours+/week is not exceptionl.

In the USA most software engineers are FLSA-exempt ("computer employee" exemption). No overtime pay regardless of hours worked. No legal maximum hours per day/week. No mandatory rest periods/breaks (federally). The US approach places the burden on the individual employee to negotiate protections or prove misclassification, while French law places the burden on the employer to comply with strict, state-enforced standa…

> 218 rest days per year (including weekends, holidays, and RTT days)

Wouldn’t that be nice, 218 rest days? It’s 218 working days.

Re: Magistral — the first reasoning model by Mistral AI

#218

I made some GGUFs for those interested in running them at https://huggingface.co/unsloth/Magistral-Small-2506-GGUF ollama run hf.co/unsloth/Magistral-Small-2506-GGUF:UD-Q4_K_XL or ./llama.cpp/llama-cli -hf unsloth/Magistral-Small-2506-GGUF:UD-Q4_K_XL --jinja --temp 0.7 --top-k -1 --top-p 0.95 -ngl 99 Please use --jinja for llama.cpp and use temperature = 0.7, top-p 0.95! Also best to increase Ollama's context length…

too much thinking

https://gist.github.com/gavi/b9985f730f5deefe49b6a28e5569d46...

Re: Magistral — the first reasoning model by Mistral AI

#219
post #98
post #67

Earlier quoted context omitted.

As an occasional user of Mistral, I find their model to give generally excellent results and pretty quickly. I think a lot of teams are now overly focused on winning the benchmarks while producing worse real results.

If so we need to fix the benchmarks.

I think there's a fundamental limit to benchmarks when it comes to real-world utility. The best option would be more like a user survey.

Re: Magistral — the first reasoning model by Mistral AI

#220

Earlier quoted context omitted.

They shouldn't be forcing people to use patented Qualcomm technology to access cellular networks either but here we are. Realistically Apple's connector adds no value and if they want to sell into markets like the EU they need to cut that kind of thing out.

> Realistically Apple's connector adds no value Like I said, usb-c is a regression from lightning in multiple ways. * Lightning is easier to plug in. * Lightning is a physically smaller connector. * USB-C is a much more mechanically complex port. Instead of a boss in a slot, you have a boss with a slot plugging into a slot in a boss. There was so much buzz around Apple no longer including a wall wort with its phones,…

I've worked with thousands of both types of cable at this point

> Lightning is easier to plug in.

according to you? neither are at all difficult

> Lightning is a physically smaller connector.

I've had lightning cables physically disassemble in the port, the size also made them somewhat delicate

> USB-C is a much more mechanically complex port.

much is a bit well, much... they're both incredibly simple mechanically — the exposed contacts made lightning more prone to damage

I've had multiple Apple devices fail because of port wear on the device. Haven't encountered this yet with usb-c

> The same logic applies to Apple forced to switch to USB, except that the costs are now multiplied.

Apple would have updated inevitably, as they did in the past — now at least they're on a standard... the long-term waste reduction is very likely worth the switch (because again, without the standard they'd have likely switched to another proprietary implementation)

Post reply on HN