Earlier quoted context omitted.
[flagged]
The bad words police are the actual authoritarian nutjobs. That is the problem.
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
81–90 of 154 posts
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#82Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#83I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…
As the old adage goes, US innovates, China imitates, EU regulates.
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#84I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…
> The kind where malicious intent is okay if the words are nice.
Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#85Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#86I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…
> "that one moderation style" we already know from current big tech platforms. > The kind where malicious intent is okay if the words are nice. Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.
For example, by positioning what you’re doing as an accessibility tool and using the right words, you can get the latest models to write incredibly powerful malware without safeguards kicking in.
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#87Earlier quoted context omitted.
Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custo…
A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather beni…
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#88Earlier quoted context omitted.
[flagged]
public ones are content moderation as above and previously llama guard, et al OCR is also another field - Mistral have a model, so do deepseek The ones I have experience with where you fine-tune smaller / faster models for business tasks like content writing, support, etc. by their nature stay private
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#89Mistral needs to abandon their Everything-stral branding. Getting kind of lame. "Shieldstral" is an awkward and bad name
On general purpose LLMs, and vibe coding, Mistral lags behind. But I find the targeted LMs much more interesting.