Record an Agent's MCP Traffic Once, Then Replay It in CI Without Servers or Credentials
Your agent test passes on your laptop and fails in CI. The laptop has a token for the issue tracker’s MCP server; CI doesn’t, and shouldn’t. Even with a token, the tracker listed three open bugs yesterday and lists four today, so the agent’s summary changes and the assertion on it breaks.
Web developers dealt with this years ago. VCR (Ruby), Polly.js, Betamax and go-vcr record real HTTP responses into a file the first time a test runs, then serve them from that file on every run after. The file is called a cassette. MCP needs the same thing: a proxy that sits between an agent and its MCP servers, writes every exchange into one cassette, and plays it back later with no server, no credentials and no network. mcprec below is an illustrative name for a tool you’d have to build.
Unit tests miss this kind of failure, which is why API testing concentrates on the boundary, and for an agent the boundary is its tool calls. A hand-written mock is the older answer, and API mocking and sandboxes covers when it’s enough. It stops scaling once an agent picks a different path through dozens of tools every run.
The HTTP cassette tools can record an HTTP MCP server in principle. They fall short in three places. Local servers still commonly use stdio, which never touches the network, so an HTTP interceptor has nothing to see. A JSON-RPC server answers everything on one path, so default matching on method and URL sees a single endpoint for every tool. And none of them knows what an interim result is, which matters below.
Keploy is closer in spirit: it records the calls a service receives along with the calls it makes to its dependencies, then replays them as tests with mocks. But it sits on the service’s side of the wire, and here the thing under test is the client. GoReplay and Speedscale replay captured traffic against a live server to test the server. This is the mirror image: the server is the missing piece, and the recording plays its part. The server-side cousin is a regression suite built from captured traffic; the single-request version is a repro file for one failed call.
The Stateless Spec Makes Matching Cheap
The 2026-07-28 MCP revision removed the initialize handshake and the Mcp-Session-Id header. Every request now carries the protocol version, client info and capabilities in _meta, and over HTTP the Mcp-Method and Mcp-Name headers name the method and the tool without anyone parsing the body. Under the old protocol a replayer had to fake a handshake and serve answers in the order that session had seen them. Now a request describes itself, so a cassette entry can be keyed on method, tool name and canonical arguments and looked up directly.
Two warnings. A stateless protocol doesn’t mean a stateless server: list_issues returns something different after create_issue, and a table keyed on arguments can’t express that (ordered matching, below, can). And _meta carries client info on every request, so a key built from the whole body breaks the day you upgrade your agent framework. The key has to skip _meta.
Recording, Matching and a Useful Miss
The recorder is a proxy with two front doors. For HTTP it’s a reverse proxy: the agent’s server URL points at it, and it copies each exchange down on the way through. For stdio it can’t see inside a process the agent spawns, so it writes a modified copy of the agent’s MCP config in which every server entry launches through mcprec, which starts the real server as a child and relays messages both ways. The transports themselves are covered in the general MCP guide.
The cassette is JSON Lines, one record per line, so it diffs cleanly. The first line is a header: format version, protocol version, the server’s name and version from server/discover, the model name and temperature if the model was recorded too, and a hash of the redaction rules. Every line after that is one exchange:
{"seq":14,"method":"tools/call","name":"search_issues","args":{"query":"label:bug state:open"},"result":{"isError":false,"content":[{"type":"text","text":"3 results: #412, #409, #377"}]},"ms":412}
Replay serves recorded results by key, and a miss is the interesting event, so it gets the most work. Print the closest recorded call with the same method and tool name, ranked by how many fields differ, and name the differing JSON Pointers (RFC 6901 paths such as /args/query).
$ mcprec record --config mcp.json --out triage.mcp -- ./run-agent.sh
recorded 31 exchanges from 2 servers, 3 secrets redacted
$ mcprec replay triage.mcp --strict -- ./run-agent.sh
MISS tools/call search_issues
asked: {"query":"label:bug state:open sort:updated"}
closest: recorded #14, {"query":"label:bug state:open"}
differs: /args/query
exit 3: 1 miss, 4 recorded exchanges never used
Three match strategies cover most cases. Exact compares canonical JSON of the arguments (sorted keys, normalized numbers). Ignore-fields drops the JSON Pointers you list first, for volatile arguments like a since timestamp. Ordered serves the Nth call for a key with the Nth recorded answer, which is how list_issues can differ before and after create_issue. VCR’s record modes include one that refuses anything unrecorded. That’s the mode CI wants.
The Hard Part Is Everything That Varies
A live model won’t make the same calls twice, even at temperature zero, because providers don’t promise identical output across runs or model updates. So replay has to say which of two different tests it’s running.
In the first, the proxy also sits on the model API’s base URL and records the model’s responses. Replay then drives the agent’s own code (prompt assembly, the tool loop, retries, output parsing) deterministically, at zero token cost. That’s a unit test of your agent’s glue code. In the second, a live model plays against the recorded world, and a call that misses gets either a hard failure (CI) or an isError: true result carrying the nearest recorded call, so the model can correct itself. That’s an eval, and it flakes like one. Keep the two apart.
Redaction happens on write, because a cassette with a live token in it is a leaked credential waiting for a push. But blanking everything to [REDACTED] breaks replay: a customer’s email shows up in one result and again in a later call’s arguments, and if two customers collapse into the same string the agent can’t tell them apart. Map each distinct secret to a stable placeholder instead (email-1, email-2), using a counter table that lives in memory during recording and is thrown away afterwards. The same value then becomes the same placeholder everywhere, so on replay the agent sees placeholders and reuses them in the calls that follow.
Multi Round-Trip Requests are the part no HTTP cassette tool has a word for. A server can answer a tools/call with an interim result, resultType: "input_required", plus request state; the client gathers the input and retries the call with that state echoed back. In the cassette that’s a chain: call, interim result carrying state S, retry carrying S, final result. Replay hands out the recorded S verbatim and matches the retry on it. Treat S as opaque bytes and keep redaction away from it.
Stream resumability is gone from the spec: a broken stream means the client re-issues the request with a new request ID, so request IDs are useless as keys. It also means a cassette can hold a failure on purpose. Record a stream that died halfway and replay it as a broken stream, to check that your agent retries instead of reporting half a result. Progress notifications are stored with their offsets and replayed in order, minus the delays. Images and audio arrive as base64 in content blocks; past a size cap, store them as hash-named blobs beside the cassette so the file stays diffable.
Then there’s staleness. Servers rename tools, tighten schemas and add result fields. Re-record on a schedule in a job that does have credentials, and diff the new cassette against the committed one. The diff, tools/list results included, tells you a tool changed before your agent trips over it, which makes the recorder change detection for your dependencies.
Scope for Version 0.1
The first version records over stdio and HTTP, writes the JSONL cassette with its header record, supports exact and ignore-field matching, prints nearest-miss diagnostics, and applies redaction rules with stable placeholders. Because request state is just part of the key, interim results replay from day one. Ordered matching and model recording come next.
It refuses two things. It won’t become a general observability tool; capturing prompts, token counts and spans across a whole agent run is a different job, and a portable agent flight recorder is the better home for it. It also won’t ship a UI for editing cassettes. They’re text, so edit them in an editor and review the diff like code.
Freeze the world first. Then you can tell whether the agent got worse.