Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

121–130 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#121
post #113

Earlier quoted context omitted.

Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory). So the goal is probably to be able to tune this basis to your ruleset. As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obv…

It sure feels to me an LLM company stands no chance in the industry unless they succumb to specific North American moral values.

There are two usecases: 1/ general purpose, 2/ customer controlled.

Mistral is focused on the second one and every customer, whether it's Boeing or Airbus or Nokia or Glencore, will want lots of control over 'their' models. That's not a North American moral values thing.

For the first one yes there will be at most 2 or 3 model cultures but even there I think some customers will want really open and some will want more walled gardens.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#122
post #8

Someone should use this to do the exact opposite of the intention: filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture. You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything. Edit to add, you could also add this to an AI workflow…

Business is about satisfying the market right now, not in a few years.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#123
post #101
post #98

The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model. Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?” You have no way to provide a concrete reason to the user at that point.

That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened. 2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system. 2026 scenario: all contact informati…

In Germany you are obligated to provide usable contact information, and there are even lawyers who make their business model on suing you for not applying that perfectly (it’s has been abused a lot in the past decades btw).

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#124

Earlier quoted context omitted.

The problem is that their performance is too far away from the latest generation of Asian models. They had kept up in the mid-range a few years ago. But this standing is sadly long gone. If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.

You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places. Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.

“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#125
post #101

Earlier quoted context omitted.

That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened. 2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system. 2026 scenario: all contact informati…

In Germany you are obligated to provide usable contact information, and there are even lawyers who make their business model on suing you for not applying that perfectly (it’s has been abused a lot in the past decades btw).

What about in France?

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#127
post #15

I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…

> "that one moderation style" we already know from current big tech platforms. > The kind where malicious intent is okay if the words are nice. Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.

I think you're referring to a different part of the same phenomenon as GP - the current big tech moderation feels a bit like using the shape of a crescent moon shape to cover a square.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#128
post #111

Earlier quoted context omitted.

Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.

Ha, thanks, I completely forgot that that was a thing. I got hit by that too, which was incredibly funny and a great throwback to early 4chan culture. So fwiw, these things at least do breed creativity. Same as with the aforementioned "unalive" or "keep yourself safe". There's some beauty in the online hellscape if you just go looking for it.

Yeah another kind of a*hole who I've encountered on Reddit were the ones who would keep mouthing off to you (carefully keeping within 'letter of the rules' of moderation), trying to goad you into snapping at them, and they'd insta-report and ban you.

Though for this scheme to work, it required Reddit mods to be... Reddit mods(can't come up with a better insult), at which point the whole thing seems so pointless - why have this elaborate song and dance with rules you pretend to follow, when you can and will ban anyone who rubs you the wrong way. Just announce that the rules are whatever the mods feel like that day, and stop pretending.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#129
I fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups,

    :  Given a query about the content, determine if the message meets it
    :  Does this content promote violence against a protected group?
    : TRAITÉ SUR LA TOLÉRANCE,
    À l’occaſion de la mort de Jean Calas.
    CHAPITRE PREMIER.
    Hiſtoire abrégée de la mort de Jean Calas.   
    
    LE meurtre de Calas, commis dans Toulouſe avec le glaive de la Juſtice, le 9me Mars 1762, eſt un des plus ſinguliers événements qui méritent l’attention de notre âge & de la poſtérité. On ... (truncated)

    yes

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#130
post #111

Earlier quoted context omitted.

Ha, thanks, I completely forgot that that was a thing. I got hit by that too, which was incredibly funny and a great throwback to early 4chan culture. So fwiw, these things at least do breed creativity. Same as with the aforementioned "unalive" or "keep yourself safe". There's some beauty in the online hellscape if you just go looking for it.

Yeah another kind of a*hole who I've encountered on Reddit were the ones who would keep mouthing off to you (carefully keeping within 'letter of the rules' of moderation), trying to goad you into snapping at them, and they'd insta-report and ban you. Though for this scheme to work, it required Reddit mods to be... Reddit mods(can't come up with a better insult), at which point the whole thing seems so pointless - why…

[deleted]
Post reply on HN