Agent Flight Recorder by Corey JacobsAgent Flight Recorder by Corey Jacobs

Agent Flight Recorder

Corey Jacobs

Corey Jacobs

Agent Flight Recorder — tool-use run evidence and failure capture for AI agents. Problem: tool-using agents fail after prompts, tool calls, state updates, or external side effects have already happened. When they do, the evidence of what went wrong is usually gone — leaving teams to guess which failure deserves to become a regression case. What I built: a local-first recorder for observable tool-using agent runs. It records model calls, tool calls, tool results, state snapshots, checkpoints, errors, and replay requests into append-only event timelines backed by SQLite-first storage, with CLI-first inspection and export, checkpoint inspection, replay tickets and plans, side-effect-aware replay helpers, and best-effort redaction at ingest. From recorded events it can export portable run bundles and generate regression-case material — including which failures should become regression cases or eval seeds. Evaluation approach: when an agent fails, the recorder preserves what the agent received, what the model returned, what tools were called, and what state resulted — so failures can be replayed, inspected, and converted into regression cases (case.json, pytest template, README) that are safe by default. The project is explicit about its boundaries: it is not model interpretability, not an enterprise security product, and does not guarantee capture of every event. Deliverables: local-first run recorder with append-only timelines; model/tool/state/checkpoint/error capture; replay tickets and side-effect-aware replay helpers; regression-case generator (case.json, pytest template, README); CLI-first inspection and export; SQLite-backed storage with ingest-time redaction; documented honest non-claims. What this demonstrates: turning agent failures into preserved, replayable, regression-ready evidence — the operational backbone of my AI Agent and Tool-Use Reliability Audit and Custom AI Eval Suite services. Evidence: https://github.com/cwwjacobs/agent-flight-recorder
Like this project

Posted Sep 30, 2026

Local-first observability for AI agents. AFR records tool/model calls, state, checkpoints, and errors so failed runs can be inspected, replayed, and tested.