timeout, read its exit code, tail its log, hold a lock, fire a curl notification. That is bash's native surface; a Python layer would add a dependency and an interpreter to the one thing that must keep working when a node is starved. The cost is paid in string discipline (set -euo pipefail, quoted paths) and in tests — 13 of them — rather than in a runtime..git/index. A worktree gives each lane its own checkout and its own branch off the same object store, so lanes are isolated without duplicating the repo. Cleanup is deferred (cc-gc.sh) and only ever deletes a branch already merged into its fan-in or main.RESULT line and a machine verify-cmd — not either alone. The RESULT: ok|fail contract is how a run declares intent cheaply; it costs one line and lets the driver classify a run that did something no external check can see (a doc rewrite, a migration plan). But a model's self-report is not evidence — so a lane only counts as done when its verify-cmd exits 0. The two catch different failures: a missing RESULT line is ambiguous (the run may have hung), a failing verify-cmd is a real fail. Conflating them is the trap the 2026-09-02 audit swarm fell into (below).task.md, out.log, usage.json, cost.log). Telemetry is a best-effort UPSERT of one row derived from those files — it never fails the run that triggered it. The database earns its keep only for the cross-run questions the files can't answer cheaply: fail rate by week, cost by model, estimator error over 200+ resolved predictions. The disk stays the source of truth; the DB is the index over it.queries.sql alongside this run's artifacts; see Links. The telemetry is live, so counts move by a few between snapshots.verify-cmd runs recorded, all 26 passing. The 13 test scripts cover the harness guardrails themselves — max-parallelism, lane timeout, verify denylist, run-id substitution, escalation cap and the sync-only / background-wait rule.task.md is written. That asymmetry is the entire reason cc-opus-gate.sh refuses an Opus run unless a reason is set. Its effect is not measurable from telemetry — refused runs never spawn, so no row is ever written for them; there is no "before" set to compare. The gate is justified by the 28%/76% split it forces a decision against, not by a measured delta.RESULT: ok|fail line disagrees with what the audit task actually meant — the 2026-09-02 swarm above is recorded fail at the swarm level despite lanes completing. Parsing a self-reported line is not the same as verifying the outcome.RESULT line is unclassifiable; the driver has to treat "no result" as its own state rather than assume success.tf-gen-20260922-1954, then tf-topics-20260922-2041 two steps later) and again on 2026-09-02 (pf-raglab-20260902-0914/-1150): the model's own log said, verbatim, that it would "wait for the background task notification" or had "scheduled a check-back" instead of polling synchronously. The harness gained an explicit sync-only rule and a BG-WAIT marker for it — the rule is appended to every task.md, but it's advisory text, not a sandbox restriction, so the same pattern can still recur.mkdir lock that refuses to start a second run on a busy path.RESULT: line in the log — and that mechanism works as designed: the driver marks them ambiguous, tags a fail_reason of BG-WAIT or NO-RESULT, notifies, and a chain halts on the branch for manual review rather than guessing.exit_code file at all and no RESULT: line anywhere in their log — the wrapper process itself died mid-run (SSH session drop, a kill, a reboot) before it ever reached its own exit trap. Checked by hand on four of them: one ends mid-turn on a blocked permission request, two never got past writing task.md (the spawn itself never started), one has a 0-byte log. None of these are flagged ok, fail, or ambiguous anywhere — they're invisible unless someone lists the directory tree by hand, which is what finding them for this page required. The mkdir isolation lock only stops a second run from starting in a busy directory; nothing currently revisits a directory a dead wrapper left behind. That's an open gap, not a solved one.task.md — advisory, so the same pattern still recurs (it did, twice on 2026-09-22). The right fix is to deny the background-notification tools at the harness level for unattended runs, so the rule can't be ignored rather than merely stated.mkdir lock stops a second run entering a busy directory but nothing revisits a directory a dead wrapper left behind — hence the 41 phantom dirs below. A periodic reaper that reconciles run-dirs against telemetry would close that gap.cost_usd and style exist in swarm.runs but the telemetry INSERT never writes them, so cost has to be reconstructed from disk by joining cost.log per run. That reconstruction is why the cost figures cover a matched subset, not all rows — it should be written at run time.cc-run.sh <run-dir> against a directory holding a task.md; chains and swarms take a plan file (lane|task|verify|model per line). The full component list, gate exit-code table, and env-var contract are in the README.local:<tag>) are planned — the code path exists as a stub today, not a working backend.Posted Sep 27, 2026
Driver fleet for unattended claude -p jobs: chains, parallel worktree swarms, per-lane verification, cost pre-flight, Postgres telemetry. Hundreds of real runs.
0
0