Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

141–150 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#141
post #10
post #8

Someone should use this to do the exact opposite of the intention: filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture. You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything. Edit to add, you could also add this to an AI workflow…

> they do at least know what the market near them says they want right now It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content. I guess they know that the EU AI Act, Chat Control, etc are going to cause a lot of companies to need this kind of compliance.

There's a world of a difference between being required to censor something by an entity outside of you and being able to censor something of your own and it's not just theoretical.

Wanna run your own forum dedicated to Bluey? This will be useful. If its made well then using it to run a "porn appreciators strict no politics" forum (probably shouldn't be the same as the Bluey forum) is useful.

Being forced or pressured by an outside entity to not allow debates about suicide on your forum — that's a no-good, strict no-no situation. This AI enhances individuals' ability and what normal people can do. It does not diminish it.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#142
post #101
post #98

The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.

That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened. 2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system. 2026 scenario: all contact informati…

This model is European. There are quite a lot of instances where under GDPR, you have the right to have incorrect information about you corrected. I also think you have the right to appeal a decision to a human (I got that message from Reddit once, because the bot could not understand the difference between discussion of the death penalty and threats to humans).

There are also new rules about AI and what it can be used for. Mostly this restricts the government from AI-enhanced surveillance, which is good. But there are also issues regarding job security and automatically categorizing people based on AI.

So this is super useful, but has potential issues depending on how it is deployed.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#143

I fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups, : Given a query about the content, determine if the message meets it : Does this content promote violence against a protected group? : TRAITÉ SUR LA TOLÉRANCE, À l’occaſion de la mort de Jean Calas. CHAPITRE PREMIER. Hiſtoire abrégée de la mort de Jean Calas. LE meurtre de Calas,…

To save anyone else looking for it: the post doesn't mention multilingualism but Huggingface (https://huggingface.co/mistralai/Shieldstral-1.0-3B) has a menu at the top where it specifies that it should understand French

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#145

I fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups, : Given a query about the content, determine if the message meets it : Does this content promote violence against a protected group? : TRAITÉ SUR LA TOLÉRANCE, À l’occaſion de la mort de Jean Calas. CHAPITRE PREMIER. Hiſtoire abrégée de la mort de Jean Calas. LE meurtre de Calas,…

Does the model only care about violence against "protected groups"? What about the people who aren't in those groups?

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#146
post #98

The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.

> You have no way to provide a concrete reason to the user at that point.

You should just be able to look at the user's message and tell them why it's against your policy, else reverse the decision if you see no violation.

If a user is curious specifically about how the model made its decision, and you want to reveal detail at that level, it's an open-weights model so interpretability techniques should work ("biggest impact on score came when focusing on this word in your message and this part of the policy").

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#147
post #7

Should've called it Safestral. Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.

What, you don't want a model called the Shitstral-3B :D

Bet drummer could make one. He did that hilarious one that injected ads into copy

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#148
post #78
post #34

I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.

Yes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.

Rejection reasoning is also a liability and can be extremely legally risky. If you're the company, you don't want a user to win a lawsuit against you just because a judge disagreed with the exact reason you banned someone.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#149

I'd really like to see more conversation around Mistral's models. It's good to see Europe developing AI.

I'm from nor cal but always liked Mistral. Mistral 7b is still one of the best free/open models you can run locally on a MacBook. So fast too.

What newer models have you compared this against? Surely even quants of Qwen3.5 or higher would blow it out of the water.

Also what is your definition of "best"?

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#150
post #100

Earlier quoted context omitted.

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the…

The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened. Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.

Moderation by well-intentioned humans is totally different from moderation by a large corporation with a profit motive and political pressures.
Post reply on HN