Skip to main content
This walkthrough builds one agent end to end: a contract reviewer that extracts governing law, indemnification and termination clauses from vendor contracts. Everything shown is from a real session.
cd into your project and run helix — you are at the prompt. Not installed yet? See Install Helix.
Driving Helix from a script or a coding agent? Send each command with helix -p "<command>" and continue the same session with helix -c -p "<next command>". Without the interactive screen Helix cannot ask its interview questions, so give the answers in the /spec prompt: the output format, what to do when a clause is missing, and where the agent runs.

Spec the contract reviewer

One sentence starts the interview:
Helix reads your project for context (a sample contract in contracts/ gets noticed and used), then interviews you — structured questions with written-out options, not a blank form:
An animated capture of the spec interview: the dashboard, the /spec command being typed, the agent reasoning about the contract-review domain, and the structured question dialog with tabs for The pain, Who uses it, Output shape and Missing clauses.

A real /spec run, trimmed and sped up — about 8 seconds of a 3-minute session.

For this agent it asked, among others: The interview ends with a validated agentspec.yaml — the portable definition of the agent. An excerpt of what this run produced:
.mutagent/specs/contract-reviewer/agentspec.yaml
The spec also carries the agent’s system prompt, its standard operating procedure (ingest → locate clauses → structure the output), and — this matters for the rest of the loop — its evaluation contract: verbatim-extraction-accuracy, missing-clause-detection, structured-json-validity. The session closes by naming the next step:

Build it

Build implements the validated spec into the target the interview chose — here, a markdown agent at .claude/agents/contract-reviewer.md — through an implement-and-verify loop, and stops at a working agent. Build has the mechanism.

Evaluate it

Evaluate scores the agent against the criteria the spec declared: did every citation match verbatim contract text, did silent clauses get flagged rather than invented, did the JSON validate. Each criterion returns pass/fail with a confidence, rolled into one verdict. Evaluate has the full pipeline.

Improve one you already have

You don’t have to start at Spec. Point Helix at an agent that already runs and has traces:
Where it flags failures, ask why:
Diagnose returns ranked remedies. Nothing is applied until you approve. Approve one in Diagnose to apply it once, or hand it to Optimize, which applies it and re-evaluates until your criteria pass.

You’re done when

  • The Helix dashboard renders (/help repaints it any time), and
  • /spec (or just asking for a spec) starts the interview, and
  • after the spec and build, .mutagent/specs/contract-reviewer/agentspec.yaml and the agent file the interview chose (here .claude/agents/contract-reviewer.md) exist in your project.
A full /evaluate needs an agent and its traces. A new agent has no traces yet: run it on a few sample contracts first (for example, ask Claude Code to use the contract-reviewer agent on them), then /evaluate contract-reviewer scores those runs.
Prefer prose to commands? Every command here has a natural-language equivalent — “I want a contract-review agent…” routes to /spec the same way. See Commands.