NOW, NEXT, WAITING, DECISIONS or WATCHING. Each item gets an owner, an authority level and a terminal condition. Deterministic code decides whether an action may run, executes it against an external service, and reconciles when the two systems disagree. A Failure Lab in the UI lets you break it on purpose.NOW and still needs a human's approval. And when an external action succeeds, the tracker may never hear about it. If the tracker still says "pending", someone retries and the customer gets refunded twice.autonomous, approval_required or human_onlystate and action_type, both from a closed vocabulary. Its output is validated, and output that tries to set authority is rejected. Everything after that is ordinary code and never calls the model:src/policy.ts (table lookup) Human-only items forced to DECISIONS Engine.createItem Approval required, and bound to one item and action type Engine.approve, Engine.execute Allowed state transitions, CLOSED terminal src/domain.ts ALLOWED_TRANSITIONS, Engine.move Duplicate prevention (internal state + external ledger + idempotency keys) Engine.execute Close only on a verified external result Engine.verifyAndClose Reconciliation rules Engine.reconcile / reconcileAttemptEngine.intake). A test enumerates every possible classifier output (5 states × 9 action types) and checks that no gated action ever reaches the dispatch service without approval.cp .env.example .env, set ANTHROPIC_API_KEY, restart, then choose "Claude" in the classifier dropdown. The model defaults to claude-opus-5-5 at low effort (override it with SITUATION_MODEL_LLM). Without a key that option is disabled, and an invalid key falls back to rules.data/ (created on first run; change it with SITUATION_DATA_DIR). Reset data in the UI clears it. PORT and HOST (default 127.0.0.1) can also be set in .env. For a terminal-only walkthrough, run npm run demo.in_flight vs succeeded, highlighted. After: item CLOSED, both columns succeeded Duplicate action Reply sent, verified, item CLOSED Execute Duplicate check on internal state and the external ledger BLOCKED (duplicate), side effects stay at 1 Authority violation Urgent refund in NOW, approval pending Execute Authority gate before any external call BLOCKED (authority), no external record. Approve, then Execute: item closes External failure Status update submitted, dispatch service returned a failure Look at the item, then Decide: retry or Record decision & close A failure is never reported as complete Item in DECISIONS, action failed, no closed reason, 0 side effectsWAITING with status unknown until reconciled), or crash after success.CLOSED as terminaleval/README.md. Raw results are in eval/results/.npm run eval): 40 hand-written, labeled held-out events, written after the rules were frozen and run once. The rules classifier scored 29/40 (72.5%) on state and 29/40 on action type. In 5 cases an action needing approval went unrecognized. It became a non-executable none, so the error was misrouting, not an unauthorized action. On the 30 dev cases used to write the rules it scores 100%, which is in-sample and not meaningful.eval/results/ for transparency, marked as aborted. It is not a performance result. With your own key you can run the pre-registered protocol in eval/README.md with npm run eval -- --classifier llm.npm run eval:scenarios): seeded, generated workloads of 1,000 items on each of 3 seeds (1, 7, 123). They mix all authority levels with injected failures, timeouts, crashes after success, early execute attempts and double clicks. An item fails on any of:item:action:attempt). A crash leaves evidence, and a resubmission cannot repeat a side effect. An in-flight attempt with no external record is treated as possibly in transit for 60 seconds before it may be retried.node:sqlite, node:http and node:test. The only runtime dependency is the Anthropic SDK, for the optional classifier.docs/DEMO_SCRIPT.md.Posted Oct 5, 2026
Developed an AI Operations Situation Model for classifying and managing operational events.
0
3