An AI Agent Flight Recorder Belongs in One Portable File, the Way HAR Did It for HTTP
An agent edits the wrong config file, then spends forty minutes trying to repair its own repair. The user files a bug with a screenshot of the last message. The maintainer asks for logs, and the logs live in four places: the provider’s usage page, the tool server’s stdout, the framework’s debug output, and a terminal scrollback that closed with the window. Nobody can say what the agent saw on turn six.
What’s missing is a file. A flight recorder for agents should write everything that crossed the boundary (prompts, model responses, tool calls, MCP traffic, timings, token counts, errors, and which call started which) into one portable file that opens the same way on any machine. The recorder is the easy part. The file format is the product, and the front ends are small programs that read it: a terminal stream in the spirit of strace, later a live view in the spirit of top.
HAR is the model to copy. The HTTP Archive format is JSON that browser devtools export, so a support engineer can say “send me a HAR” and get one. It recorded what crossed the wire, with timings, and said nothing about the server’s insides, which kept it small and universal. It also taught a lesson about risk: HAR files carry cookies and auth headers, so people learned to scrub them before attaching them to a ticket. An agent recording has the same problem with worse contents, because prompts hold source code and customer text.
The existing tools aren’t aimed at this. LangSmith, Langfuse, Arize Phoenix, Helicone and OpenLLMetry are all built around sending traces somewhere, and a trace inside a hosted product is a link that needs a login. OpenTelemetry’s GenAI semantic conventions, still marked as in development, are settling attribute names, which fixes the vocabulary but gives you nothing to attach to a ticket. Use their names wherever they exist; a rival vocabulary would kill adoption on day one. The same idea at smaller scale is packaging one failed API request into a file; an agent run is a few hundred requests with a tree on top.
What Goes in the File
The unit is a span: an ID, a parent ID, a kind (turn, model call, tool call, MCP call, sub-agent), start and end times, a status, and for model calls the model ID, provider-reported token counts and a cost estimate. Parent links give you the tree: a turn contains a model call, which asked for two tools, one of which started a sub-agent. A flat log loses exactly that.
Payloads go in a blob table keyed by SHA-256, hashed per message rather than per request. Turn 30’s request repeats messages 1 through 29, so storing whole requests grows quadratically; storing each message once and recording a request as an ordered list of hashes costs almost nothing. The same trick keeps a 20,000-token system prompt, and the tool definitions, to one copy each. It also answers the question bug reports keep asking, “why did it forget what I said on turn 3?”: compare the hash list for turn 12 with turn 11 and the dropped message is right there. Context window management as a database covers how a framework decides what to drop.
For the container, one SQLite file with four or five tables beats a zip of JSONL, and crashes are the reason. An agent that dies mid-run should still leave a readable file, and a zip keeps its index at the end, so a recording cut short leaves an archive many tools refuse to open. SQLite in WAL mode with synchronous=NORMAL survives the process dying; a power cut can lose the last few commits, which is a fair price. Anyone with sqlite3 can query it, and a short export --jsonl covers people who prefer grep. The spec is a few pages plus sample files for testing independent readers. Here’s a session, with a made-up command name.
$ agenttrace --out bug.agentlog -- ./run-agent.sh "fix the failing billing test"
0.000 t1 model big-model in 18204 out 212 1.9s est $0.06
1.914 t1 tool shell pytest -x tests/billing exit 1 3.2s
5.140 t2 model big-model in 19107 out 388 2.4s est $0.07
7.560 t2 mcp tickets search q="billing flaky" ok 410ms 8.1KB
8.002 t2 tool edit_file billing/retry.py ok 4ms
8.011 t3 model big-model 429 retry-after 2s retried
10.140 t3 model big-model in 20840 out 96 1.1s est $0.07
$ agenttrace show bug.agentlog --errors
Every line there is a row in the file, so the stream and the file can’t disagree. Whoever receives bug.agentlog runs show and sees what the author saw, filtered by --tools, --tokens, --errors or --cost.
Capturing Without Patching Every SDK
Model traffic is the easy half. Point the SDK’s base URL at a local proxy (most SDKs read an environment variable for it or take a constructor argument) and the recorder sees every request and response, reassembling streamed ones into a single record while passing the bytes through untouched.
MCP is the second half. Over HTTP the same proxy works, and the 2026-07-28 revision helps: every request carries its protocol version and client info in _meta, and the Mcp-Method and Mcp-Name headers label a tools/call for search before the body is parsed. Local servers mostly use stdio, so the recorder has to launch the server itself and relay JSON-RPC between agent and child, writing both directions into the file. A cassette for replaying MCP traffic in CI is tuned for matching; a flight recorder is tuned for reading.
A proxy can’t see in-process tools, file edits or subprocesses. The wrapper form, agenttrace -- <command>, follows child processes and watches file and network activity where the OS allows, the way strace -f follows forks. Where it doesn’t (restricted containers, some desktop systems), the file should say “not observed”, because a gap that looks like silence misleads. In-process events go through a local socket: a framework emits a span with a dozen lines in its agent loop, and that cheapness is how a format spreads.
One layer sits below all of this. VPN Works runs an agent in a Linux network namespace whose only exit is its own command, checking each connection against a short policy and writing it to a record (how that works). It sees the socket a tool opened that the model never mentioned; the flight recorder sees the prompt that led there. Together they answer what the agent thought, called and connected to.
The File Is Sensitive by Nature
A recording copies everything the agent was shown: source code, tickets, customer messages, the occasional API key pasted into a chat. Redact on write, before anything reaches disk. Auth headers are never stored. Patterns for known key formats and a user-supplied list catch the rest, and each redacted value becomes a placeholder derived from a keyed hash with a random per-file key, so the same secret shows as the same placeholder everywhere. A reader can see a token was reused without being able to guess it. A shape-only mode keeps names, timings, sizes, token counts and hashes, drops the blobs, and is often enough to find a loop or a slow tool.
Ordering is its own trap. Several processes, maybe several machines, each with a clock. Trust one: the recorder’s. Timings measured at the proxy are what the agent experienced, and durations a server reports about itself are stored as separate claims. Each event gets a sequence number on arrival. Timestamps can disagree; a sequence can’t.
A recording also has honest limits. It shows what crossed the boundary, and it can’t show what the model considered and dropped. If the provider returns reasoning text it’s stored like any other output; if not, the file has a gap, and the viewer should say so. An OTLP exporter belongs in the first release (agenttrace export bug.agentlog --otlp http://localhost:4318), so existing backends can ingest the file. The span model is the one covered in API observability, with turns where that piece has services.
Version 0.1
The first version does proxy capture for two model API shapes, MCP over HTTP and stdio, the file with its spec and sample files, show and tail in the strace style, shape-only mode, redaction rules and the OTLP exporter. The stream is the hook: one command, no account, no backend. The top view comes next, as a reader that polls the file while the recorder writes. It refuses a hosted backend, evaluations, prompt management and any UI beyond the terminal.
The payoff comes from unattended runs. A scheduled agent that goes wrong at 3 a.m. leaves nobody to ask, and cron for agents can tell you a run was killed, not why.
Get the file right. Somebody else will build the dashboard.