A text moderator that classifies prompts and responses using a defined hazard taxonomy; this profile examines Llama Guard 3–8B.
Specialized models help check prompts, answers, and actions. Explore where they fit, what to evaluate, and how practitioners might use them.
A text moderator that classifies prompts and responses using a defined hazard taxonomy; this profile examines Llama Guard 3–8B.
A small classifier for attempts to override instructions, distinct from a general harmful-content moderator.
Google’s moderation family includes separate text and image models. Pick the modality before comparing capabilities.
A family of judges for harmful content, groundedness and tool-call risks, with custom judging criteria in newer releases.
Qwen3Guard offers complete-message and streaming moderation, with safe, controversial and unsafe labels.
A reasoning-based classifier that evaluates content against a supplied policy, making policy quality part of the evaluation.
NVIDIA’s Llama 3.1 Nemotron Safety Guard 8B v3 moderates text interactions and reports unsafe categories.
Ai2’s 7B moderator evaluates prompt harmfulness, response harmfulness and whether an answer is a refusal.
Rules, permissions, sandboxes, and monitoring play different roles.