Live data from Hacker News

Magistral — the first reasoning model by Mistral AI

mistral.ai

341–350 of 444 posts

Re: Magistral — the first reasoning model by Mistral AI

#341
post #339

One immediate observation I have about this model is that it seems to do a better job of filtering out or toning down some ideological disinformation that other models regurgitate from activist controlled Wikipedia articles, at least for a few I've checked. Previously you had to write your own sanity-check prompts to get the model to do extra up-front work to validate the logical and historical accuracy of things bef…

What’s an example prompt and sanitized prompt you use to evaluate?

Re: Magistral — the first reasoning model by Mistral AI

#342
We just tested magistral-medium as a replacement for o4-mini in a user-facing feature that relies on JSON generation, where speed is critical. Depending on the complexity of the JSON, o4-mini runs ranged from 50 to 70 seconds. In our initial tests, Mistral returned results in 34–37 seconds. The output quality was slightly lower but still remain acceptable for us. We’ll continue testing, but the early results are promising. I'm glad to see Mistral prioritizing speed over raw power, there’s definitely a need for that.

Re: Magistral — the first reasoning model by Mistral AI

#343

Earlier quoted context omitted.

Indeed, and with the technology plateau-ing, being 6-12 months late with less debt is just long term thinking. Also, Europe being in the race is a big deal for consumers.

>with the technology plateau-ing People were claiming that since year 2022. Where's the plateau?

Could it be that at least for the "lowest" fruits, most amazing things that can one can hope to obtain from scraping the whole web and throw it at some computation training was already achieved? Maybe AGI simply can not be obtained without some relevant additional probes sent in the wild to feed its learning loops?

Re: Magistral — the first reasoning model by Mistral AI

#344
post #223
post #218

Earlier quoted context omitted.

too much thinking https://gist.github.com/gavi/b9985f730f5deefe49b6a28e5569d46...

My impression from running the first R1 release locally was that it also does too much thinking.

Magistral Small seems wayyy too heavy-handed with its RL to me:

\boxed{Hey! How can I help you today?}

They clearly rewarded the \boxed{...} formatting during their RL training, since it makes it easier to naively extract answers to math problems and thus verify them. But Magistral uses it for pretty much everything, even when it's inappropriate (in my own testing as well).

It also forgets to unless you use their special system prompt reminding it to.

Honestly a little disappointing. It obviously benchmarks well, but it seems a little overcooked on non-benchmark usage.

Re: Magistral — the first reasoning model by Mistral AI

#345

Earlier quoted context omitted.

The cookie banners are corps trying to circumvent the rights and protections. If they actually went by the spirit of the protections, the cookie banners wouldn't be needed. Your ire is misdirected.

Are you sure? The ePrivacy Directive requires a (GDPR-level) consent for just placing the cookie, unless it's strictly necessary for the provision of the “service”. The way EU regulators interpret this, even web analytics falls outside the necessity exception and therefore requires consent. So as long as the user doesn't and/or is not able to automatically signal consent (or non-consent) eg via general browser-level…

As you said yourself, analytics are not necessary.

It's corpos trying to invade our privacy.

Re: Magistral — the first reasoning model by Mistral AI

#346

Earlier quoted context omitted.

This is not true. Government workers or factory workers can limit to 35h (with some salary loss or days off loss), but else than that (especially in tech) it is very competitive and working 50 hours+/week is not exceptionl.

In the USA most software engineers are FLSA-exempt ("computer employee" exemption). No overtime pay regardless of hours worked. No legal maximum hours per day/week. No mandatory rest periods/breaks (federally). The US approach places the burden on the individual employee to negotiate protections or prove misclassification, while French law places the burden on the employer to comply with strict, state-enforced standa…

I think there is theory and there is real life. As tech worker, in 20 years career, in private sector, I have always been on forfait jours, working more than 10h/day on average, during many years weekend included. I never got paid extra hours. So I get what you say about the perception and the law. The French law is protective (i.e if I can prove that in a court I'll get my extra hours paid for sure but my career would end. Period.

Re: Magistral — the first reasoning model by Mistral AI

#347
post #30

Is the number of em-dashes in this marketing copy indicative of the kind of output that the model produces? If so, might want to tone it down a bit.

This meme that humans don’t use em dashes needs to die.

It’s an extremely useful tool in writing and I’ve been using it for decades.

Re: Magistral — the first reasoning model by Mistral AI

#348
post #119

Here are my notes on trying this out locally via Ollama and via their API (and the llm-mistral plugin) too: https://simonwillison.net/2025/Jun/10/magistral/

> I guess this means the reasoning traces are fully visible and not redacted in any way - interesting to see Mistral trying to turn that into a feature that's attractive to the business clients they are most interested in appealing to.

but then someone found that, at least for distilled models,

> correct traces do not necessarily imply that the model outputs the correct final solution. Similarly, we find a low correlation between correct final solutions and intermediate trace correctness

https://arxiv.org/pdf/2505.13792

ie. the conclusion doesn't necessarily follow from the reasoning. So is there still value in seeing the reasoning? There may be useful information in the reasoning, but I'm not sure it can be interpreted by humans as a typical human chain of reasoning, maybe it should be interpreted more as a loud multi-party discussion on the relevant subject which may have informed the conclusion but not necessarily lead to it.

OTOH, considering the effects of automation fatigue vs human oversight, I guess it's unlikely anyone will ever look at the reasoning in practice, except to summarily verify that it's there and tick the boxes on some form.

Re: Magistral — the first reasoning model by Mistral AI

#349

Earlier quoted context omitted.

Because we do not have a complete understanding of human neurons. How are we supposed to accurately model something we cannot directly observe?

Do you also complain when someone says "Half-life 2 has great water-physics" with "Don't call it physics, we still don't understand all the physical laws of the universe, and also they use limited-precision floating-point, so it's not water-physics, it's just a bunch of math"? Like, we've agreed that "water-physics" and "cloth physics" in 3d graphics refers to a mathematical approximation of something we don't actual…

That’s comparing apples to oranges. Nobody is going to be making a real cruise ship based on game water physics simulations.

In such a task, better water simulations are used. We have those, because we can directly observe the behavior of water under different conditions. It’s okay because the people doing it are explicitly aware that they are using simulation.

AI will get used in real decisions affecting other people, and the people doing those decisions will be influenced by the terminology we choose to use.

Re: Magistral — the first reasoning model by Mistral AI

#350

Earlier quoted context omitted.

It does not do any thinking. It is a statistical model, just like the rest of them.

These kind of comments are the equivalent of going to dog owners' forums, analyzing word choices in every post and warning the dog owners about the dangers of anthropomorphizing their pets, an effort as accurate as it is boorish and ineffectual.

[flagged]
Post reply on HN