Skip to main content
Helix scores your agent against real traces — the runs your agent already produces in development or production. It reads what you already have; you don’t add instrumentation for Helix.

Supported sources

Langfuse

Read traces from a Langfuse project.

OpenTelemetry

Any OTel-compatible trace export.

Local JSONL

A file of runs on disk — no service required.

Session transcripts

Your Claude Code or Codex session transcripts, as-is.
If you log to something that isn’t listed, the local JSONL path is the escape hatch — export a batch of runs to a file and point Helix at it.

What a trace needs to be useful

A trace is one run of your agent: the input, what the agent did, and the output. The more of the run that’s captured — tool calls, intermediate steps, the final result — the more precisely the evaluator can judge and the diagnostics can root-cause. A bare input/output pair still works; a full trajectory works better.
No traces yet? You can still start. Run Spec and Build to create the agent; Evaluate and Diagnose become useful the moment it starts logging runs.

Where your data goes

Helix reads traces where they already are (session files on your machine, OTLP/JSON files, or your Langfuse project) and does not copy them to Mutagent. Traces reach Mutagent only when you send them: Claude Code sessions through the hooks mutagent hooks install sets up, and the runs of Helix Cloud sandboxes. When the evaluator judges, it calls the model provider you signed in with, on your own credentials, so there is no inference bill from us.