Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

101–110 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#101
post #98

The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.

That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened.

2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system.

2026 scenario: all contact information has been scrubbed from the site. Users can click “chat” and a chatbot will apologize for their dissatisfaction and offer no option to escalate. No one will ever hear anything about the complaint, so there’s no need to explain the failure. User can either accept this or can get f**ked because all competitors operate the same way.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#102
post #78
post #34

I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.

Yes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.

I think IRL in the “rejection” case they don’t want to tell the user exactly why, since the user may be malicious and use it to try to evade the block. And for use in moderating UGC, well, most platforms don’t take seriously the idea that they need to answer to their users. Only their advertisers.

In the case of wondering why a bad thing got through, well, I think that’s why they just set these to the most pro-censorship level they can, to make that highly unlikely.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#103
I'm liking the trend of companies are releasing smaller, focused models instead of trying to make one model do everything. A dedicated moderation model is much easier to reason about than hideden safety logic inside a general-purpose model which might not have had much training in that aspect at all

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#104
post #98

The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.

Yeah, you do. You go and review it manually if the user comes back and says that.

The reality is that most users don’t ask because they know they violated the rule.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#105
Folks should check out https://zentropi.ai and their latest model, CoPE-B-A4B: https://huggingface.co/zentropi-ai/cope-b-a4b

Policy adaptive models really are the coolest things these days.

Also, check out https://roost.tools for even more open safety tooling!

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#106
post #101
post #98

The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.

That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened. 2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system. 2026 scenario: all contact informati…

Not all platforms want to operate like mainstream social media though

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#107
post #51

Earlier quoted context omitted.

Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]:…

do you have a source for their GB300 count?

The 13,800 number seem like it would be quoted from their annoucement of the $830 million funding round they had in March.

Here's CNBC's article on it: https://www.cnbc.com/2026/03/30/mistral-ai-paris-data-center...

Not sure if they would have received the full number yet, but it's been a few months so they certainly could have. Bit of a moot point when the comparison was against Poolside's Laguna which isn't really "general" SOTA but SOTA-for-the-size, and Mistral is clearly capable of training 700B or 120B models that are that when released considering they have done that... A 2-3T model is probably possible with the GPUs they have but they would need to spend most of their resources on it, and it's not clear why they would want to.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#109
post #100
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the…

The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened.

Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#110
post #100
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the…

Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.
Post reply on HN