Introducing Shieldstral.
Introducing Shieldstral.
Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.
Selection record
AdmittedSum of both 139 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:发布AI多模态安全分类器模型
- Why it was chosen
- A 3B classifier with natural language policies replaceable at inference time demonstrates the usable scope of small models for moderation.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Mistral AI · mistral.ai