Audience
AI platform teams that need customizable, multimodal content moderation without retraining a separate safety model for every policy
About Shieldstral
Shieldstral is a 3B open-weights, policy-adaptive multimodal safety classifier designed to evaluate text, images, and text-plus-image content using policies defined at inference time. Instead of relying on a fixed taxonomy of harm categories, it frames moderation as a binary question-answering task: users provide an instruction describing the evaluation context and strictness, a yes-or-no safety question, and the content to judge. The model reads the “yes” and “no” logits and converts them into a continuous, calibrated safety score, allowing applications to threshold or rank results by confidence rather than depend on a single discrete label. This formulation unifies prompt classification, response moderation, refusal detection, toxicity detection, and multimodal safety in one interface, while letting teams adapt policies without retraining the model. Shieldstral can evaluate prompts, responses, prompt-response pairs, images, and images with accompanying text.