One agent — a contract reviewer — taken from a sentence to a validated spec, a build, and a first verdict.
This walkthrough builds one agent end to end: a contract reviewer that extracts governing law,
indemnification and termination clauses from vendor contracts. Everything shown is from a real
session.
cd into your project and run helix — you are at the prompt. Not installed yet? See
Install Helix.
Driving Helix from a script or a coding agent? Send each command with helix -p "<command>"
and continue the same session with helix -c -p "<next command>". Without the interactive screen
Helix cannot ask its interview questions, so give the answers in the /spec prompt: the output
format, what to do when a clause is missing, and where the agent runs.
/spec a contract-review agent that extracts governing law, indemnification and termination clauses from vendor contracts
Helix reads your project for context (a sample contract in contracts/ gets noticed and used), then
interviews you — structured questions with written-out options, not a blank form:
A real /spec run, trimmed and sped up — about 8 seconds of a 3-minute session.
For this agent it asked, among others:
Question
The answer taken
How should the agent structure the extracted clauses?
Structured JSON with exact clause text, section citations, and key-term summaries
When a contract is silent on a topic, how should it react?
Flag as missing/silent with a risk notice — never fabricate
Where will the agent run?
A Claude Code markdown agent (.claude/agents/contract-reviewer.md)
The interview ends with a validated agentspec.yaml — the portable definition of the agent. An
excerpt of what this run produced:
.mutagent/specs/contract-reviewer/agentspec.yaml
metadata: id: contract-reviewer name: Contract Reviewerspec: intent: outcomes: - "Extract governing law jurisdiction, indemnification coverage/triggers, and termination terms into validated structured JSON." - "Provide exact section citations and verbatim clause excerpts for legal verification." - "Explicitly flag missing or silent clauses with risk notices." constraints: - "Never fabricate or hallucinate contract language." - "Must provide exact section references and verbatim text excerpts for every clause." nonGoals: - "Negotiating contract language or redlining terms." - "Providing definitive legal advice or enforceable legal opinions."
The spec also carries the agent’s system prompt, its standard operating procedure (ingest →
locate clauses → structure the output), and — this matters for the rest of the loop — its
evaluation contract: verbatim-extraction-accuracy, missing-clause-detection,
structured-json-validity. The session closes by naming the next step:
### Next Lifecycle StepTo turn this Definition into executable code, run:/build
Build implements the validated spec into the target the interview chose — here, a markdown agent at
.claude/agents/contract-reviewer.md — through an implement-and-verify loop, and stops at a working
agent. Build has the mechanism.
Evaluate scores the agent against the criteria the spec declared: did every citation match verbatim
contract text, did silent clauses get flagged rather than invented, did the JSON validate. Each
criterion returns pass/fail with a confidence, rolled into one verdict.
Evaluate has the full pipeline.
You don’t have to start at Spec. Point Helix at an agent that already runs and has traces:
evaluate contract-reviewer against last month's reviews
Where it flags failures, ask why:
why is contract-reviewer missing indemnification clauses in order forms?
Diagnose returns ranked remedies. Nothing is applied until you approve. Approve one in Diagnose to
apply it once, or hand it to Optimize, which applies it and re-evaluates until your criteria pass.
The Helix dashboard renders (/help repaints it any time), and
/spec (or just asking for a spec) starts the interview, and
after the spec and build, .mutagent/specs/contract-reviewer/agentspec.yaml and the agent file
the interview chose (here .claude/agents/contract-reviewer.md) exist in your project.
A full /evaluate needs an agent and its traces. A new agent has no traces yet: run it on a few
sample contracts first (for example, ask Claude Code to use the contract-reviewer agent on them),
then /evaluate contract-reviewer scores those runs.
Prefer prose to commands? Every command here has a natural-language equivalent — “I want a
contract-review agent…” routes to /spec the same way. See
Commands.