Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

91–100 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#91
post #34

I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.

OpenAI's moderation API is multi-modal and free with no strings attached in a way that truly boggles the mind.

I've put easily over a billion requests (>$100,000 by typical moderation API pricing) through it over the last few years for $0.

I think it's a severely underappreciated offering, but I also don't bother pushing it too hard because who knows when the party will end lol. Strikes me as something that's only stuck around because no one's abusing it.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#92
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

if you pull stuff like that off, and it gets flagged, it is obvious you are trying to game the system. Same with a system like this. It could be used in addition to human moderators, where the human moderator have to meta-moderate the AI's work. This could also be used as training exercise for new moderators. Then, the good and experienced moderators have more time to spend on edge cases, complex cases, fine-tune the AI, etc. In other words, it is able to do the most boring things to you.

And EU regulate, I mean as a counter example: we are not in panic about a nipple. When I was in a large museum in Paris, multiple women were breast feeding their infant. And why not? Kid's gotta eat. I'll refrain from insulting any world leaders, too easy, but you know many examples are available there regarding censorship.

Finally, it can take that BS argument away of 'oh we don't have manpower to moderate'. That is a low blow, too, by large commercial entities who could, you know hire and train? However, even a small company with not much money to burn could -in theory- win here.

I'd give this model a chance, if not only cause I've been impressed by Mistral past years. Yes, Le Chat / Vibe probably lags behind, but something like Voxtral (real-time and transcribe) is neat, and efficient.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#94
Crazy that it's a small lab becoming the frontier in term of moderation models, instead of Meta which is pouring dozens of billions into LLMs.

Meta would really benefit from work done on this front, however their model Llama Guards are quite lagging compared to the competition.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#96
post #68
post #66

Earlier quoted context omitted.

Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custo…

A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather beni…

> In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment.

In French as well. But it’s obviously not what’s meant here.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#98
The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model.

Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?”

You have no way to provide a concrete reason to the user at that point.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#99
post #51

Earlier quoted context omitted.

It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.

Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]:…

do you have a source for their GB300 count?

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#100
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the thinly-coded language, so why are we censoring everything in the first place?
Post reply on HN