Mistral's Shieldstral: 3B open-weights model for multimodal moderation
Mistral has released Shieldstral, a 3B parameter open-weights multimodal model designed for content moderation. It uses a policy-adaptive question-answering approach, allowing developers to define safety rules at inference time without needing to retrain the model.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in