Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

81–90 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#83
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

> Kinda like cultural imperialism but with an ethical spin.

As the old adage goes, US innovates, China imitates, EU regulates.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#84
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

> "that one moderation style" we already know from current big tech platforms.

> The kind where malicious intent is okay if the words are nice.

Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#86
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

> "that one moderation style" we already know from current big tech platforms. > The kind where malicious intent is okay if the words are nice. Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.

I thought OP was referring to AI platforms.

For example, by positioning what you’re doing as an accessibility tool and using the right words, you can get the latest models to write incredibly powerful malware without safeguards kicking in.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#87
post #68
post #66

Earlier quoted context omitted.

Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custo…

A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather beni…

Discriminative has a meaning in machine learning that I think is relevant here. There are "generative" models like LLMs that are learning joint probabilities P(X, Y) and "discriminative" models like logistic regression that learn conditional probabilities P(Y | X)

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#88
post #80
post #77

Earlier quoted context omitted.

[flagged]

public ones are content moderation as above and previously llama guard, et al OCR is also another field - Mistral have a model, so do deepseek The ones I have experience with where you fine-tune smaller / faster models for business tasks like content writing, support, etc. by their nature stay private

Voxtral [1], maybe Robostral? I just saw on their blog they released OCR 4. Their is usually informative [2]

[1] https://mistral.ai/news/voxtral-transcribe-2/

[2] https://mistral.ai/news/

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#89
post #4

Mistral needs to abandon their Everything-stral branding. Getting kind of lame. "Shieldstral" is an awkward and bad name

Argue about taste. At least it is more original than OpenAI (haha, 'open') who started as non-profit and then pulled an Infantino. The Le Chat logo is also cool retro :)

On general purpose LLMs, and vibe coding, Mistral lags behind. But I find the targeted LMs much more interesting.

Post reply on HN