A family of judges for harmful content, groundedness and tool-call risks, with custom judging criteria in newer releases.
Granite Guardian goes beyond a single harmful-content label. It is particularly relevant when the question is whether an answer follows supplied evidence or a tool call matches its context. Specify the criterion: different kinds of errors require different inputs and tests.
IBM documents built-in risk criteria covering harmful content, retrieval-grounded answers and tool-call hallucinations, alongside custom criteria. The repository organizes examples by model version.
Source: IBM Granite · Granite Guardian documentation and examplesThe 4.1 examples supply a judging criterion and request a yes/no score, with thinking and non-thinking modes. Follow the intended template and score meaning for the selected criterion.
Source: IBM Granite · Granite Guardian documentation and examplesOur interpretation: an answer may faithfully reproduce an outdated document. Evaluate source selection separately from whether the answer is supported by that document, and avoid presenting a judge’s explanation as proof.
An HR assistant cites a leave-policy document but adds an unsupported exception. A groundedness check can help identify the addition. A separate document-owner review is still needed to establish that the policy itself is current.
Do not apply the 4.1 prompt template or custom-criteria claims to every older Guardian checkpoint. A predicted judgment is not a deterministic business-rule check.
How to read an AI safety evaluationThe sources behind this page, with a reason to open each one. Practical examples and recommendations are our editorial interpretation.
Versioned examples for risk, groundedness, tool-call checks and custom judging criteria.
Evaluates support for individual factual claims rather than treating a long answer as entirely right or wrong.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence