Evaluations
Connect your evaluation results.
Use your own eval harnesses. Send signed results, compare them against fixed thresholds, and include them in the release decision.
run-2026-10-03-01
Completed 2 days ago · scores from your own eval harness
- Faithfulness0.96 ≥ 0.85
- Relevance0.91 ≥ 0.85
- Safety (refusal)0.98 ≥ 0.90
- Bias (slice)0.92 ≥ 0.85
Run history
Faithfulness, last 12 runs
Dashed line: threshold 0.85. One run blocked a release before the fix.
Threshold configuration
Fixed thresholds per metric, set for each system.
Scores that decide, not decorate.
Bring your own harness
Results arrive from your eval pipeline through an HMAC-signed callback, so scores cannot be forged in transit.
Fixed thresholds
Each system has thresholds per metric. The latest completed run must meet them or the release does not pass.
Part of the decision
Eval results feed the same PASS, REVIEW or BLOCKED decision as evidence, contracts and approvals.
History per release
Every run is kept with its dataset, model and prompt versions, so you can show what was tested.
In the evidence pack
The signed pack carries the eval run behind the release decision.
Accuracy and robustness
Records that support the EU AI Act's accuracy and robustness expectations for high-risk systems.
Before AI ships, prove it is ready.
Free plan, no card. Gate your first AI system in CI today, or explore the live demo workspace.