Earlier quoted context omitted.
Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custo…
A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather beni…
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
71–80 of 154 posts
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#72Earlier quoted context omitted.
It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.
Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]:…
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#73I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…
Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custo…
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#74Earlier quoted context omitted.
Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training. The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to. [1]:…
Isn't poolside a completely different company from Mistral?
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#75Is it honest about religious texts? Can I throw at it religious texts and it'll honestly tell me whether the text promotes physical violence or not?
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#76Someone should use this to do the exact opposite of the intention: filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture. You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything. Edit to add, you could also add this to an AI workflow…
> they do at least know what the market near them says they want right now It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content. I guess they know that the EU AI Act, Chat Control, etc are going to cause a lot of companies to need this kind of compliance.
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#77I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms. The kind where malicious intent is okay if the words are nice. ___ Or, rephrased: How big is the space in which you can tune this model without retraining. Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _…
> which seems to be mistrals whole thing They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche. Before the datacenter deals their revenue was higher than xAI's There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost. Mistral, Microsoft model releases and T…
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#78I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#79> ... a single yes/no question, e.g. "Does this content promote physical violence?" Is it honest about religious texts? Can I throw at it religious texts and it'll honestly tell me whether the text promotes physical violence or not?
Re: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
#80Earlier quoted context omitted.
> which seems to be mistrals whole thing They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche. Before the datacenter deals their revenue was higher than xAI's There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost. Mistral, Microsoft model releases and T…
[flagged]
OCR is also another field - Mistral have a model, so do deepseek
The ones I have experience with where you fine-tune smaller / faster models for business tasks like content writing, support, etc. by their nature stay private