What is Hedera Harness?
Hedera Harness is a complete agentic loop for Hedera dApps. Give it a written spec and a repo, and it goes the whole way on its own: reads the spec, implements it, installs, builds, runs your tests, drives the running app in a real browser, reads state back off the network, repairs what failed, and tries again. Then it tells you whether the spec was met and hands you the evidence behind that answer. A spec is the only thing it needs from you. Everything betweenharness run and the verdict happens without you in the loop — walk away, or run it in CI.
And the harness decides whether a run passed, not the agent. It runs your install, build, and test commands itself. The running app is judged by a second agent that is never allowed to see the source. Evidence for any claimed failure has to exist on disk before the verdict is accepted. An agent reporting on its own work is the failure mode this exists to remove.
Use it to:
- Ship a feature into a repo you already have, written as a spec instead of a prompt
- Walk away from the loop — implement, test, judge, and repair all run unattended
- Get a verdict you can defend, with screenshots, network reads, and command output behind it
- Run the whole thing in CI with
--json
Quick start
Install, answer four questions, run your first spec
How a run works
Doctor, derive, generate, test, evaluate — and the repair loop
Writing a spec
The one rule that decides whether a run can pass
GitHub
Source, issues, and contributing guide
What you point it at
Any Hedera dApp repo with apackage.json, a git history, and a clean working tree. The harness reads your own install, build, test, and serve commands out of harness.yaml rather than imposing a layout, so it works on a project that already exists and was never built with it in mind.
No project yet? Create one with Scaffold HBAR, then hand the harness a spec for the first feature you want on top of it.
Quick start
Prerequisites
- Node.js ≥ 20 — nodejs.org
- Git, and a repo with at least one commit and a clean working tree
- A
package.jsonat the repo root - Claude Code credentials — either a subscription login or
ANTHROPIC_API_KEYin your environment - Chromium, for the evaluation stage:
doctor stage checks every one of these before a run starts, and tells you exactly which is missing.
Specs that write to the network need a funded Hedera testnet account. Export
HEDERA_OPERATOR_ID and HEDERA_OPERATOR_KEY in the shell that runs the harness; the harness passes the key to the judging agent and scrubs it from every artifact it writes. Read-only specs need neither.Install
Set up and run
harness init asks four questions, with defaults read from your package.json. Use it when you already know what you want built.
harness wizard payment-flow does the same job with an agent: it reads the project, works the answers out itself, then interviews you one question at a time and drafts specs/payment-flow.md. Use it when you do not yet have a spec.
harness.yaml and commit it, because it describes the project rather than any one run. See the reference for its fields and for every command flag.
How a run works
Only the first stage can stop to ask you anything, and only with
--review. Everything after it runs to a verdict on its own.
Each attempt is committed to a harness/<timestamp> branch. When the run ends you are back on the branch you started from, and the work is there to inspect, cherry-pick, or throw away.
If a run stops early — you interrupt it, or it hits a budget — every completed attempt is already committed. harness run --spec <path> --continue picks up from that branch rather than from yours, starting by repairing what the last attempt failed on. A run that passed is continuable in the same way, which is how a second spec builds on the first.
How it decides
A passing run is only worth something if the verdict can be trusted. Five rules make that true:- The judge cannot see your code. The evaluating agent runs in a directory holding only the spec, with the repo denied at the sandbox boundary. An agent that can read the implementation ends up grading the implementation instead of the app.
- A verdict must show its work. Every claimed failure cites evidence, and the harness confirms those files exist on disk before it accepts the verdict.
- The checklist is written before anything is built. Derive reads the spec against an app that does not exist yet, so the app cannot lower the bar it is later measured against.
- The harness keeps the books. The judge reads chain state before it touches the app and again afterwards, and the verdict is not accepted until every checklist item is accounted for: both readings, whether it held, and the evidence that shows it.
- A verdict is never re-rolled. If the judge answers, that answer stands. Only the absence of an answer earns a second look.
Reading a finished run
report answers what live output cannot: what each attempt did, what the run cost, and why you should believe the verdict.
--full to include both agents’ complete tool feeds, so you can read what the judge actually checked rather than taking its word. Every artifact a run produces is kept under .harness/runs/<timestamp>/, which the harness excludes from git for you.
Related tools
Scaffold HBAR
Bootstrap the Hedera dApp you point the harness at
Hedera Skills
The Hedera knowledge the building agent loads during a run