Live data from Hacker News

Magistral — the first reasoning model by Mistral AI

mistral.ai

431–440 of 444 posts

Re: Magistral — the first reasoning model by Mistral AI

#431
post #408
post #83

Earlier quoted context omitted.

For me this is more important than quality. I love fast responses, feels more futuristic.

What are you using LLMs for?

Mostly "How is $term in English", what is $thing, review this message for clarity, clean up this data, parse this screenshot, and coding.

Re: Magistral — the first reasoning model by Mistral AI

#432
post #428

Earlier quoted context omitted.

When each reply gets longer and longer it's a sign there's less and less chance of finding common ground. I'll try to make this one brief: - It takes two minutes to write an email pointing out an egregious comment or bad actor. - People are flagging and complaining about your comments because they are breaking the guidelines, nothing more, nothing less. Maybe you're not aware of it due to a cultural disconnect. If th…

I wrote a long comment before just to explain you my thought process in detail so you know that I'm here for good faith debates, not to break rules, but OK, I'll make it short for you now: You still haven't explained why my comment on the Android topic I was replying to is in bad faith but the on I replied to isn't, when I explained you in detail why it is, by the same yardstick you used to judge my comment? You keep…

If you're asking why this comment [1] wasn't flagged: no users flagged it, because it doesn't break the guidelines. It's not inflammatory. It doesn't set off a flamewar. It doesn't fulminate. It just raises a question, for people to respond to. It got some upvotes and some downvotes and some replies, most of which were fine. Yours was the only one that was inflammatory and broke the guidelines. Not egregiously, but enough, given your recent patterns.

> You keep saying it's not you who decides what's right and wrong, that it's the users who decide based on the rules by flagging

I think this is a misunderstanding. We moderators don't (can't) make judgements about accuracy or truthfulness of comments. All we can do is determine if a comment breaks the guidelines. Comments should only be flagged by users if they break the guidelines. Our enforcement of the guidelines is independent of the accuracy of the comment's content or its ideology.

If a comment of yours is flagged for any reason other than guidelines breaches, you're within your rights to protest. But given your conduct even in this subthread with me, in which your comments continue to be full of guidelines breaches, it seems you're not able to gauge whether your comments are within the guidelines or not.

It's going to keep being a problem if you're not able to correct that.

[1] https://news.ycombinator.com/item?id=44240808

Re: Magistral — the first reasoning model by Mistral AI

#433

Earlier quoted context omitted.

> Countries like Greece, Italy, Spain, Portugal PIGS, really? Some of the top growing EU economies right now, which have turned their deficit around, show the future of a slowly stagnating Europe?

A 200B economy growing 2% is the future of the EU? Yes that is the point I am making.

How much is an economy supposed to grow?

Re: Magistral — the first reasoning model by Mistral AI

#434
post #30

Is the number of em-dashes in this marketing copy indicative of the kind of output that the model produces? If so, might want to tone it down a bit.

This meme that humans don’t use em dashes needs to die. It’s an extremely useful tool in writing and I’ve been using it for decades.

Same here. The em dash has been maybe my favorite punctuation since at least the early 2000s. All the em dash output from LLMs looks really natural to me.

Re: Magistral — the first reasoning model by Mistral AI

#435
post #353

Earlier quoted context omitted.

What’s an example prompt and sanitized prompt you use to evaluate?

Not going to leak my tests, but here's how you can create your own. - Think up a topic that's interesting to you, yet maybe controversial. - Look up primary sources and empirical information about it. - Then look at a relevant Wikipedia article about it to see if the way the Wikipedia article frames it is honestly and faithfully justified by the primary sources and empirical data about it. If the article seems to hav…

I don't get it.

If I think that Rabbits and Hares are classified by Wikipedia incorrectly and Ideological Wikipedia Editors are hiding the truth with Disinformation, why would I give the model any credit if it tells the correct answer only if I develop a custom 22 point mammalian biology reasoning checklist that leads it to the Real Truth about Rabbits and Hares?

It certainly doesn't inspire confidence that any other particular question would be answered correctly?

Re: Magistral — the first reasoning model by Mistral AI

#436
post #431
post #408

Earlier quoted context omitted.

What are you using LLMs for?

Mostly "How is $term in English", what is $thing, review this message for clarity, clean up this data, parse this screenshot, and coding.

I see! So for these, you tend to find the accuracy "good enough" on the faster-but-less-accurate models.

I generally find the same thing for simple definitions/translations and other "chat" tasks. I'm a little bit surprised that you also find it so for coding, but otherwise I think I get it.

Re: Magistral — the first reasoning model by Mistral AI

#437

Earlier quoted context omitted.

Do we know that with certainty? Do we actually? Because my understanding is that how "thinking" works is actually still a total mystery. How is it we no for certain that the basis for the analog electric-potential-based computing done by neurons is not based on statistical prediction? Do we have actual evidence of that, or are you just doing "statistical token prediction" yourself?

You’re reversing the burden of proof in a similar manner as religious people often do. Absence of evidence is not evidence of absence, and so on.

I'm not reversing it lol. You're the one making a claim, the burden of evidence is on you.

Absence of evidence is not evidence of absence, but it is still absence of evidence. Making a claim without any is more religious that not. After all, we know humans can't be descended from monkeys!

Re: Magistral — the first reasoning model by Mistral AI

#438

Earlier quoted context omitted.

> https://arxiv.org/abs/2503.09211 I am thoroughly unimpressed by this paper. It sets up a vague strawman definition of "thinking" that I'm not aware of anyone using (and makes no claim it applies to humans) and then knocks down the strawman. It also leans way too heavy on determinism - For one thing, we have no way of knowing if human brains are deterministic (until we solve whether reality itself is). For another,…

Whodathunkit, some people are so infatuated with their simulacra that they choose to go tooth and nail in defense of the simulation. My point was congruent with the argument that LLMs are not humans or possess human-like thinking and reasoning, and you have conveniently demonstrated that.

> My point was congruent with the argument that LLMs are not humans or possess human-like thinking and reasoning, and you have conveniently demonstrated that.

I mean, they are obviously not humans, that is trivially true, yes.

I don't know what I said makes you believe I demonstrated that they do not possess human-like thinking and reasoning, though, considering I've mostly pointed out ways they seem similar to humans. Can you articulate your point there?

Re: Magistral — the first reasoning model by Mistral AI

#439

Earlier quoted context omitted.

With how amazing the first R1 model was and how little compute they needed to create it, I'm really wondering how the new R1 model isn't beating o3 and 2.5 Pro on every single benchmark. Magistral Small is only 24B and scores 70.7% on AIME2024 while the 32B distill of R1 scores 72.6%. And with majority voting @64 the Magistral Small manages 83.3%, which is better than the full R1. Since I can run a 24B model on a reg…

It's not better than full R1; Mistral is using misleading benchmarks. The latest version of R1, R1-0528, is much better: 91.4% on AIME2024 pass@1. Mistral uses the original R1 release from January in their comparisons, presumably because it makes their numbers look more competitive. That being said, it's still very impressive for a 24B. I'm really wondering how the new R1 model isn't beating o3 and 2.5 Pro on every s…

Mistral isn’t using misleading benchmarks. I linked to DeepSeek’s own benchmark results that DeepSeek created. I couldn’t find anything newer.

Can you link me to the benchmark you found?

Re: Magistral — the first reasoning model by Mistral AI

#440
post #425

Earlier quoted context omitted.

> Why do you think that is? Is it not a reflection of the userbase bias? Where comments get flagged not based on rules but based on which political side they are targeting? That comment was a breach of the guidelines but it almost always takes longer than half an hour for a comment to be flagged, and for us to see it, especially on a thread that's over a day old that barely anybody is looking at anymore. You could fl…

>The fact that you didn't Mate, I don't have time to flag all comments that I find inflammatory, especially when I flagged many comments in the past and nothing happened to them, so what's the point? I flagged this one after and the comment was still there. So why are you throwing the blame on me? Why didn't you remove that comment after I pointed it out? > Again, when you see this, email us. Mate please, be serious,…

Your comment has a lot of misunderstandings, but here's the biggest one:

> Then by that yardstick, isn't the comment I was replying too also in bad faith, just like I pointed out initially?

No, it isn't. First, it's completely invalid to say "by that yardstick" - you're comparing completely different things that have no bearing on each other. Second, no, there's zero evidence that the comment you replied to was in bad faith.

The definition of a bad faith argument is one that in inauthentic, and that the argument-maker doesn't actually believe in themselves. Factually, there's no evidence to support your accusation that that comment by @butlike was in bad faith - they didn't make any self-contradictory statements in their comment, nor did they post a single other comment in that whole thread, nor did they say that would indicate that they were acting anything but genuinely.

And, factually, you were breaking the guidelines by assuming bad faith about https://news.ycombinator.com/item?id=44240808.

When challenged on it in https://news.ycombinator.com/item?id=44244578, you said:

> I assumed good faith, but then I used critical thinking and decided it's in bad faith then explained why. You don't need to agree with me on this.

This has three falsehoods in it. First, you did not assume good faith - you assumed bad faith, because there was zero evidence to support the idea that it was in bad faith. Second, you didn't use critical thinking - again, because there was no evidence to support that belief. Third, you did not explain why the comment was in bad faith - you explained why you disagreed with it, indicating that you don't understand the difference between disagreeing with someone's statements, and them being in bad faith (which is further reinforced in the above when you say "And I replied that's in bad faith since the alternative to Android 16 shitty UI is not going back to Android 2 to make Android 16 look good" - no, that's literally not what "bad faith" means).

Finally, more generally, beyond the falsehoods and fallacies that you've been making, you're also acting extremely abrasively, in ways that break the guidelines and that antagonize other users.

The theme of HN is intellectual curiosity. The way that you've been acting is the exact opposite of that.

Post reply on HN