Skip to content
Assurance OS

Evaluations

Connect your evaluation results.

Use your own eval harnesses. Send signed results, compare them against fixed thresholds, and include them in the release decision.

Sample data
Latest runTrendsCompareThresholds

run-2026-10-03-01

Completed 2 days ago · scores from your own eval harness

View all runs
  • Faithfulness0.96 ≥ 0.85
  • Relevance0.91 ≥ 0.85
  • Safety (refusal)0.98 ≥ 0.90
  • Bias (slice)0.92 ≥ 0.85

Run history

Faithfulness, last 12 runs

Dashed line: threshold 0.85. One run blocked a release before the fix.

Threshold configuration

Fixed thresholds per metric, set for each system.

View thresholds

Scores that decide, not decorate.

  • Bring your own harness

    Results arrive from your eval pipeline through an HMAC-signed callback, so scores cannot be forged in transit.

  • Fixed thresholds

    Each system has thresholds per metric. The latest completed run must meet them or the release does not pass.

  • Part of the decision

    Eval results feed the same PASS, REVIEW or BLOCKED decision as evidence, contracts and approvals.

  • History per release

    Every run is kept with its dataset, model and prompt versions, so you can show what was tested.

  • In the evidence pack

    The signed pack carries the eval run behind the release decision.

  • Accuracy and robustness

    Records that support the EU AI Act's accuracy and robustness expectations for high-risk systems.

Before AI ships, prove it is ready.

Free plan, no card. Gate your first AI system in CI today, or explore the live demo workspace.