Stage 3 of the agentic development lifecycle (ADLC) · run with
/evaluate.What goes in, what you type, what comes out
Code checks run before the model judge. A criterion that can be decided by code (a field is
present, a value matches, the output parses as JSON or matches a schema) is checked deterministically
first, with no model call, and gives the same answer on every run. Only the criteria that need
judgement go to the model judge.
Evaluate turns your traces into a pass/fail verdict. It doesn’t ask you to write the eval: it
derives the criteria from your own runs, judges each one, and checks the judge against your labels
before relying on it.
Run it
Ask for a verdict against your real runs:- Score the Deep Research agent’s answers for citation accuracy and flag which criteria fail.
- Check how often the Refund Processing agent approves disputes that should have been escalated.
What it does
1
Derive criteria
Reads your traces and derives binary, actionable success criteria from the runs that worked and
the runs that didn’t — no criteria list to hand-write.
2
Run the code checks
Criteria that code can decide (presence, exact values, valid JSON, schema conformance) are
checked first, deterministically. They cost no model call and give the same result every run.
3
Judge
For what remains, one critique-before-verdict judge per run scores every criterion in parallel
with the other runs: pass, fail or indeterminate, with a confidence. Each judge reasons before
it rules and cites the trace values it used, so every verdict arrives with its rationale. A
second check re-examines failures on gating criteria (criteria with critical or high severity;
severity is set per criterion, and high is the default) before they count.
4
Validate the judge
Measures each judge against your own labels: runs you marked pass or fail yourself, in the
review page Evaluate opens in your browser (Pass, Fail or Defer per run). With fewer than 60
labelled runs, a judge is marked
unvalidated: its verdicts still count, and the report shows
which ones have not been checked against your judgement yet.5
Gate
Rolls the per-criterion results into one verdict on the whole run:
fail, incomplete or
pass, in that order of precedence. A failed critical or high-severity criterion sinks it; one
the judge could not decide makes it incomplete rather than a quiet pass.Judge only — it never fixes
The evaluator decides; it does not change your agent. Failures route to Diagnose, which proposes the fixes. Separating the grader from the fixer is what keeps the verdict independent of the change it triggers.Use it on its own
You don’t have to run the whole loop to use this stage. Once Helix is installed, point Evaluate at your traces and it derives criteria from them and scores them, with each judge checked against your labels. No spec or build pass is required first.Next: Diagnose
Root-cause what evaluation flagged.