Skip to main content
Stage 3 of the agentic development lifecycle (ADLC) · run with /evaluate.

What goes in, what you type, what comes out

Code checks run before the model judge. A criterion that can be decided by code (a field is present, a value matches, the output parses as JSON or matches a schema) is checked deterministically first, with no model call, and gives the same answer on every run. Only the criteria that need judgement go to the model judge. Evaluate turns your traces into a pass/fail verdict. It doesn’t ask you to write the eval: it derives the criteria from your own runs, judges each one, and checks the judge against your labels before relying on it.

Run it

Ask for a verdict against your real runs:
Other ways people ask:
  • Score the Deep Research agent’s answers for citation accuracy and flag which criteria fail.
  • Check how often the Refund Processing agent approves disputes that should have been escalated.

What it does

1

Derive criteria

Reads your traces and derives binary, actionable success criteria from the runs that worked and the runs that didn’t — no criteria list to hand-write.
2

Run the code checks

Criteria that code can decide (presence, exact values, valid JSON, schema conformance) are checked first, deterministically. They cost no model call and give the same result every run.
3

Judge

For what remains, one critique-before-verdict judge per run scores every criterion in parallel with the other runs: pass, fail or indeterminate, with a confidence. Each judge reasons before it rules and cites the trace values it used, so every verdict arrives with its rationale. A second check re-examines failures on gating criteria (criteria with critical or high severity; severity is set per criterion, and high is the default) before they count.
4

Validate the judge

Measures each judge against your own labels: runs you marked pass or fail yourself, in the review page Evaluate opens in your browser (Pass, Fail or Defer per run). With fewer than 60 labelled runs, a judge is marked unvalidated: its verdicts still count, and the report shows which ones have not been checked against your judgement yet.
5

Gate

Rolls the per-criterion results into one verdict on the whole run: fail, incomplete or pass, in that order of precedence. A failed critical or high-severity criterion sinks it; one the judge could not decide makes it incomplete rather than a quiet pass.

Judge only — it never fixes

The evaluator decides; it does not change your agent. Failures route to Diagnose, which proposes the fixes. Separating the grader from the fixer is what keeps the verdict independent of the change it triggers.

Use it on its own

You don’t have to run the whole loop to use this stage. Once Helix is installed, point Evaluate at your traces and it derives criteria from them and scores them, with each judge checked against your labels. No spec or build pass is required first.

Next: Diagnose

Root-cause what evaluation flagged.