Live data from Hacker News

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

mistral.ai

21–30 of 154 posts

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#22
post #16

Earlier quoted context omitted.

The problem is that their performance is too far away from the latest generation of Asian models. They had kept up in the mid-range a few years ago. But this standing is sadly long gone. If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.

By this logic the Chinese should have just given up and let the American AI companies have the market because they were so far behind. I'm sure Europe has the capability to distill other people's frontier models to catch up if they wish to do so.

Distilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#24
post #12

Earlier quoted context omitted.

Some of the best social media is heavily moderate. This includes HN and r/credible defense . With a Quiet transparent and cheap LLM I imagine a social media website where you can have good discussion about everything around the world it would be a game changer and on my to-do list.

> Some of the best social media is heavily moderate Heavily moderated by humans with discretion. Not AI chat bots following a rules engine.

No, they are moderated by arbitrary company moderation policies not humans with independent thoughts.

AI can do exactly that.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#25
post #12

Earlier quoted context omitted.

Some of the best social media is heavily moderate. This includes HN and r/credible defense . With a Quiet transparent and cheap LLM I imagine a social media website where you can have good discussion about everything around the world it would be a game changer and on my to-do list.

> Some of the best social media is heavily moderate Heavily moderated by humans with discretion. Not AI chat bots following a rules engine.

I'd rather be censored by an AI than a reddit mod.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#26
post #22
post #16

Earlier quoted context omitted.

By this logic the Chinese should have just given up and let the American AI companies have the market because they were so far behind. I'm sure Europe has the capability to distill other people's frontier models to catch up if they wish to do so.

Distilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.

US judges have already rules that output of an LLM can't be copyrighted so not sure what would prevent Chinese companies to use said output for distillation purposes.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#27
post #4

Mistral needs to abandon their Everything-stral branding. Getting kind of lame. "Shieldstral" is an awkward and bad name

Was this one the last stral for you? The stral the broke the camel's back?

The shortest stral has been pulled for you

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#28
post #7

Should've called it Safestral. Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.

It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.

Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation

#30
post #26
post #22

Earlier quoted context omitted.

Distilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.

US judges have already rules that output of an LLM can't be copyrighted so not sure what would prevent Chinese companies to use said output for distillation purposes.

> already rules that output of an LLM can't be copyrighted

Mind sharing such cases? I'm not aware of any so far. There's the one with images, but that's commonly miss-understood, that case was ruled on a technicality (i.e. copyright needs to be attributed to a person, not a model)

Post reply on HN