Earlier quoted context omitted.
The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened. Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.
Which is the most insidious kind of automated moderation. In an algorithmic feed like TikTok, strongly downranking content and removing it are close to the same. And by doing it silently with no way to easily get feedback on what you did wrong people self-censor both the things you are moderating and the things people imagine you would like to moderate
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
151–154 of 154 posts
That's the point
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#152[dead]
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#153Earlier quoted context omitted.
What, you don't want a model called the Shitstral-3B :D
Bet drummer could make one. He did that hilarious one that injected ads into copy
Haha yeah, rivermind. A drummer tune that flips it around entirely and only lets through unsafe stuff would be pretty hilarious ngl.
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#154Earlier quoted context omitted.
could the long s `ſ` be throwing the model off?
Interesting hypothesis. Replacing "ſ" with "s" did not change the output. I think the simple explanation is the likely one (the reason I deliberately chose this specific benchmark): the model isn't intelligent enough to figure out use/mention distinctions. It understands Voltaire is discussing injustice, violence, tolerance; but it doesn't understand which side he's on.
thanks for testing it!
I guess that would make sense for such a small model to be missing the subtlety