nightshift — parallel agent orchestrator by Mark Mandryknightshift — parallel agent orchestrator by Mark Mandryk

nightshift — parallel agent orchestrator

Mark Mandryk

Mark Mandryk

A driver fleet for running unattended coding agents — sequential chains and parallel worktree swarms — with a machine verify gate on every unit of work, cost pre-flight before spawn, and full run telemetry in Postgres.
BashClaude Codegit worktrees PostgreSQLsystemdMIT

Problem

An unattended agent run has no human to catch it when it drifts. Left alone it will idle on a notification that never comes, report success on a job that failed its own acceptance test, or quietly burn budget with no ceiling. Scaling from one run to a fleet multiplies every one of those failure modes. The work needed a harness, not more prompting.

Architecture

.box{fill:oklch(18% .035 355);stroke:oklch(30% .05 353);rx:6} .t{fill:oklch(96% .012 352);font:13px sans-serif} .s{fill:oklch(68% .035 352);font:11px sans-serif} .ln{stroke:oklch(62% .27 300);stroke-width:1.5;fill:none;marker-end:url(#a)} .lime{fill:oklch(87% .22 132)} plan filelane|task|verify|model lane 1 · worktreehaiku · verify-cmd lane 2 · worktreehaiku · verify-cmd lane N · worktreeMAXPAR ceiling fan-inSonnet Postgrestelemetry

Decisions and trade-offs

bash, not Python. The whole job is process control — spawn a subprocess under timeout, read its exit code, tail its log, hold a lock, fire a curl notification. That is bash's native surface; a Python layer would add a dependency and an interpreter to the one thing that must keep working when a node is starved. The cost is paid in string discipline (set -euo pipefail, quoted paths) and in tests — 13 of them — rather than in a runtime.
A git worktree per lane, not a lock or branch-switch. Parallel lanes each mutate a working copy. Sharing one copy behind a lock serializes the thing that was supposed to be parallel; switching branches in place races on .git/index. A worktree gives each lane its own checkout and its own branch off the same object store, so lanes are isolated without duplicating the repo. Cleanup is deferred (cc-gc.sh) and only ever deletes a branch already merged into its fan-in or main.
A self-reported RESULT line and a machine verify-cmd — not either alone. The RESULT: ok|fail contract is how a run declares intent cheaply; it costs one line and lets the driver classify a run that did something no external check can see (a doc rewrite, a migration plan). But a model's self-report is not evidence — so a lane only counts as done when its verify-cmd exits 0. The two catch different failures: a missing RESULT line is ambiguous (the run may have hung), a failing verify-cmd is a real fail. Conflating them is the trap the 2026-09-02 audit swarm fell into (below).
Postgres, not just the files on disk. Every run already leaves files (task.md, out.log, usage.json, cost.log). Telemetry is a best-effort UPSERT of one row derived from those files — it never fails the run that triggered it. The database earns its keep only for the cross-run questions the files can't answer cheaply: fail rate by week, cost by model, estimator error over 200+ resolved predictions. The disk stays the source of truth; the DB is the index over it.

Numbers

Snapshot 2026-09-26. Every figure below has its source query in queries.sql alongside this run's artifacts; see Links. The telemetry is live, so counts move by a few between snapshots.
A representative parallel job — a design-compliance audit run on 2026-09-02 — fanned out to 11 lanes in two batches under a parallelism ceiling of two, and metered ~22.0M input and ~150.7K output tokens (11 lanes plus a Sonnet fan-in).
The pre-flight estimator is directionally useful but not precise: of 309 predictions, 231 have since been resolved against an actual run, with a median absolute error of ~81% (skewed by outliers to a 353% mean) — it beats a seed guess but nobody should budget a 5h/weekly cap against it to the dollar.
Machine verification is thin on real data so far — 26 verify-cmd runs recorded, all 26 passing. The 13 test scripts cover the harness guardrails themselves — max-parallelism, lane timeout, verify denylist, run-id substitution, escalation cap and the sync-only / background-wait rule.

What the guardrails target

Honest answer first: this harness has no clean before/after. Guardrails were added continuously, not at one cutover, and the 2026-09-22 consolidation into this repo postdates almost all of the telemetry — so a fail-rate drop can't be pinned on any one rule. What the data does support:
Cost, not prompt-tuning, is the lever — and the gate targets it. Opus is 120 of 431 metered run+lane rows (28%) but carries $3,493 of $4,571 spend (76%). ~95% of a run's token volume is cache-read of a growing transcript, so model choice dominates cost far more than how tightly task.md is written. That asymmetry is the entire reason cc-opus-gate.sh refuses an Opus run unless a reason is set. Its effect is not measurable from telemetry — refused runs never spawn, so no row is ever written for them; there is no "before" set to compare. The gate is justified by the 28%/76% split it forces a decision against, not by a measured delta.
The $2,795 overpay figure has a counterfactual, stated. It is the sum of what those Opus runs would have cost at Sonnet tier (Sonnet is exactly 5× cheaper on every price component). It assumes the runs would have produced equivalent output at Sonnet — which is precisely the thing that isn't guaranteed. That uncertainty is why the gate demands a written reason rather than banning Opus outright.
Ambiguous detection fires uniformly, not post-hoc. The 11 exit-code-4 runs are spread across the window — 6 in the week of 08-24, 3 in 08-31, 2 in 09-21 — i.e. the mechanism catches the pattern whenever it occurs, it wasn't bolted on after one bad week. The one visible fail spike (week of 08-24: ~22 fails of 132 run+lane rows, ≈17%, dropping to low single digits after) sits right after the first release and its 12-finding hardening pass; I won't over-claim which change drove the decline.

Failure modes found in production

RESULT vs. audit semantics. A run can exit clean yet its RESULT: ok|fail line disagrees with what the audit task actually meant — the 2026-09-02 swarm above is recorded fail at the swarm level despite lanes completing. Parsing a self-reported line is not the same as verifying the outcome.
Missing RESULT block. A run that ends without emitting its final RESULT line is unclassifiable; the driver has to treat "no result" as its own state rather than assume success.
Background-wait deadlock. A headless run that backgrounds a step and then waits for a completion notification idles until timeout — there is no channel to deliver that notification. Caught live twice on 2026-09-22 in the same chain (tf-gen-20260922-1954, then tf-topics-20260922-2041 two steps later) and again on 2026-09-02 (pf-raglab-20260902-0914/-1150): the model's own log said, verbatim, that it would "wait for the background task notification" or had "scheduled a check-back" instead of polling synchronously. The harness gained an explicit sync-only rule and a BG-WAIT marker for it — the rule is appended to every task.md, but it's advisory text, not a sandbox restriction, so the same pattern can still recur.
Two runs into one output tree. Two parallel runs writing the same web directory (a staging vs. production collision) corrupt each other — resolved with a mkdir lock that refuses to start a second run on a busy path.
Divergent driver copies. The driver scripts had drifted into three separate copies across a repo and two skills; consolidating them into this single-source repo was the reason nightshift was extracted as a standalone product.

Hung and phantom runs

History is less clean, and the gap is real, not just unlucky timing. 11 runs since 2026-08-22 exit with code 4 — clean process exit, no RESULT: line in the log — and that mechanism works as designed: the driver marks them ambiguous, tags a fail_reason of BG-WAIT or NO-RESULT, notifies, and a chain halts on the branch for manual review rather than guessing.
What isn't caught: 41 run directories between 2026-08-05 and 2026-09-20 have no exit_code file at all and no RESULT: line anywhere in their log — the wrapper process itself died mid-run (SSH session drop, a kill, a reboot) before it ever reached its own exit trap. Checked by hand on four of them: one ends mid-turn on a blocked permission request, two never got past writing task.md (the spawn itself never started), one has a 0-byte log. None of these are flagged ok, fail, or ambiguous anywhere — they're invisible unless someone lists the directory tree by hand, which is what finding them for this page required. The mkdir isolation lock only stops a second run from starting in a busy directory; nothing currently revisits a directory a dead wrapper left behind. That's an open gap, not a solved one.

What I'd do differently

Make the sync-only rule a sandbox restriction, not advice. Today the background-wait deadlock is prevented by appended text in every task.md — advisory, so the same pattern still recurs (it did, twice on 2026-09-22). The right fix is to deny the background-notification tools at the harness level for unattended runs, so the rule can't be ignored rather than merely stated.
Sweep dead run-dirs instead of only blocking new ones. The mkdir lock stops a second run entering a busy directory but nothing revisits a directory a dead wrapper left behind — hence the 41 phantom dirs below. A periodic reaper that reconciles run-dirs against telemetry would close that gap.
Populate the columns the schema already has. cost_usd and style exist in swarm.runs but the telemetry INSERT never writes them, so cost has to be reconstructed from disk by joining cost.log per run. That reconstruction is why the cost figures cover a matched subset, not all rows — it should be written at run time.
Treat the estimator as a range, not a point. At ~81% median absolute error it's honest to publish, but a driver that budgets against it should consume the IQR band, not the median — and I'd gate spawns on the upper bound.

Running it

Every state path — run directories, worktrees, DB credentials, notification hook — is overridable by env var, with defaults pointing at the maintainer's node; nothing is hard-wired to this VPS. A single run is cc-run.sh <run-dir> against a directory holding a task.md; chains and swarms take a plan file (lane|task|verify|model per line). The full component list, gate exit-code table, and env-var contract are in the README.

What's next

Tier-0 local-model lanes (local:<tag>) are planned — the code path exists as a stub today, not a working backend.
Idempotent run IDs so a re-run with an existing RESULT is a no-op.

Links

Like this project

Posted Sep 27, 2026

Driver fleet for unattended claude -p jobs: chains, parallel worktree swarms, per-lane verification, cost pre-flight, Postgres telemetry. Hundreds of real runs.