Skip to main content

Commands

Run them from inside the target repository.

Flags

The watchable output and --json are two renderers over the same event stream, so they cannot drift apart. --json carries raw values — "costUsd": 0.48 rather than ~$0.48 of tokens — which is the reason the two exist separately.

harness.yaml

Written by init or wizard, and committed, because it describes the project rather than any one run.
Each value is a string, or { run, cwd } when the command must run somewhere other than the repo root:
The comments are the agent’s reasoning from wizard, kept so the next person to read the file knows why this command and not the obvious-looking one. Script names mislead often enough to be worth the sentence: in a Scaffold HBAR project the root has no build, test, or dev at all, and inside packages/nextjs, start is a dev server while serve runs the production build. Edit any line by hand — the harness will not overwrite it.

Environment variables

Machine-level settings are environment variables, deliberately kept out of harness.yaml, which is shared with everyone who clones the repo.
.env files can never be staged by the harness, so a key cannot reach git history through a run — the enforceable half of secret handling, since the harness owns the commit. Exporting operator credentials in your shell, rather than writing them into the repo, is the other half, and that one is yours.

What a run leaves behind

result.json’s history holds one entry per attempt, with the failures it produced and the id each hashes to — which is how the report can tell you a failure is the same one as last time rather than a new one. .harness/ is added to .git/info/exclude, so run artifacts never appear in your diffs and never need a .gitignore entry.

Branches and interrupting

Each attempt is committed to a harness/<timestamp> branch, and the run puts you back on the branch you started from when it ends. If a run stops before it finishes — you interrupt it, or it runs out of budget — every attempt it completed is already committed to that branch. harness run --spec <path> --continue picks up from there: it branches from that work rather than from yours, and the first attempt starts by repairing what the last one failed on instead of reading the spec against code that already exists. A run that passed is equally continuable. That is how a second spec builds on the first.

Limits

Every stage is bounded by a clock — generation, evaluation, each command, and the dev server becoming ready. A breach ends that stage with a reason rather than hanging the run, and a run stops starting new attempts after four hours. There is deliberately no spend limit by default. What a run is worth is yours to decide, and a figure chosen here would only ever be wrong for somebody. Pass --max-spend when you want one.