Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

111–120 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#111
post #100

Earlier quoted context omitted.

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the…

Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.

Ha, thanks, I completely forgot that that was a thing. I got hit by that too, which was incredibly funny and a great throwback to early 4chan culture.

So fwiw, these things at least do breed creativity. Same as with the aforementioned "unalive" or "keep yourself safe".

There's some beauty in the online hellscape if you just go looking for it.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#112
post #51

Earlier quoted context omitted.

It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.

Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]:…

From my experience their capabilities are extremely narrow and generally perform terribly when faced with issues outside comparatively narrow training data.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#113
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory).

So the goal is probably to be able to tune this basis to your ruleset.

As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obvious differences: hate speech and sex). Mistral is an european company with a french basis, so I doubt they didn't plan for that (otherwise they're complete morons, which I don't think they are).

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#114
post #100

Earlier quoted context omitted.

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the…

The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened. Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.

Which is the most insidious kind of automated moderation. In an algorithmic feed like TikTok, strongly downranking content and removing it are close to the same. And by doing it silently with no way to easily get feedback on what you did wrong people self-censor both the things you are moderating and the things people imagine you would like to moderate

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#115

I'd really like to see more conversation around Mistral's models. It's good to see Europe developing AI.

The problem is that their performance is too far away from the latest generation of Asian models. They had kept up in the mid-range a few years ago. But this standing is sadly long gone. If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.

You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places.

Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#117
post #42

Earlier quoted context omitted.

I am not sure how reliable it is in the real world. Also, in terms of liability, I don’t know how effective it would be to satisfy various regulations compared to a human moderator team.

I hear ya, but one could set different operating thresholds: auto-approve low-risk posts, hold ambiguous posts for review, and automatically reject very high-confidence violations. So HITL for sure, but MUCH less H in the L.

Human-somewhere-nearish-the-loop!

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#118
post #113
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory). So the goal is probably to be able to tune this basis to your ruleset. As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obv…

It sure feels to me an LLM company stands no chance in the industry unless they succumb to specific North American moral values.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#119
post #100

Earlier quoted context omitted.

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the…

Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.

When did that feature get taken offline?

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#120
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

> "that one moderation style" we already know from current big tech platforms. > The kind where malicious intent is okay if the words are nice. Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.

It's the one where you get censored for using words like "kill" or "shit", but if you say "I'm going to come to your house and unalive you" nothing happens because you didn't trip the word filter. When the company gets in trouble for this, to fix it, the company starts removing all comments that include the word "house".
Post reply on HN