Mistral's Shieldstral: 3B open-weights model for multimodal moderation

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

Mistral AI releases Shieldstral, a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.

The same content can be fine for a cybersecurity research tool and harmful on a mental-health platform.

More from this day

2026-08-04