Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

151–154 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#151

Earlier quoted context omitted.

The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened. Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.

Which is the most insidious kind of automated moderation. In an algorithmic feed like TikTok, strongly downranking content and removing it are close to the same. And by doing it silently with no way to easily get feedback on what you did wrong people self-censor both the things you are moderating and the things people imagine you would like to moderate

That's the point

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#153

Earlier quoted context omitted.

What, you don't want a model called the Shitstral-3B :D

Bet drummer could make one. He did that hilarious one that injected ads into copy

Haha yeah, rivermind. A drummer tune that flips it around entirely and only lets through unsafe stuff would be pretty hilarious ngl.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#154

Earlier quoted context omitted.

could the long s `ſ` be throwing the model off?

Interesting hypothesis. Replacing "ſ" with "s" did not change the output. I think the simple explanation is the likely one (the reason I deliberately chose this specific benchmark): the model isn't intelligent enough to figure out use/mention distinctions. It understands Voltaire is discussing injustice, violence, tolerance; but it doesn't understand which side he's on.

thanks for testing it!

I guess that would make sense for such a small model to be missing the subtlety

Post reply on HN