Shieldstral 1.0 (3B)
Description
Shieldstral is a 3B open-weight multimodal safety classifier from Mistral. It frames content moderation as policy-adaptive yes/no question answering: plain-language policies are supplied at inference time, and the model returns a calibrated safety score from a single forward pass for text, image, or text+image content. Released under Apache 2.0. Aggregate self-reported average F1: text-safety 0.849, multimodal 0.838 (HF mistralai/Shieldstral-1.0-3B / arXiv 2607.25857); per-bench WildGuard/HarmBench/ToxicChat/BeaverTails/Aegis/VLGuard ids are not in the catalog.
Capability Radar
Science is estimated from LLM Stats science scores or reasoning when no dedicated science benchmarks are available.
Rankings
No ranking data available
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Safety
AA Evaluation Indices
(Artificial Analysis)No AA evaluation data available
LLM Stats Category Scores
(LLM Stats (zeroeval))Pricing
No pricing data available
Speed
No speed data available
Provider Price Ranking
No provider data available