Package a Failed API Request Into One File Anyone Can Replay Locally
A customer’s checkout returns a 500. Support pastes the request ID into the ticket, and the engineer on call finds the log line: KeyError: 'tax_region' in the pricing module. They send the same request locally and get a 200. Of course they do. Their database has no customer with a null tax region, the feature flag that routes to the new tax engine is off in development, the rates service answers differently today, and the clock is a day later. The bug is a function of all of that, and the ticket contains none of it.
Bug reports die in the gap between “it failed in production” and “it fails on my machine”. What’s missing is a bundle: the sanitized request, the config and flag values the server saw, the database rows the request read, the outbound calls it made and what came back, the code version, and one command that replays the lot. The version worth building has a file extension, a manifest and a spec, so frameworks can write it and test runners can read it. HAR did that for browser traffic: every browser exports one, and plenty of tools import it. Server-side request capture has nothing with that kind of reach yet.
What Already Exists
rr, from Mozilla, records and replays a process’s execution, thread scheduling included. It’s the strongest tool in this family and built for debugging sessions, not for sitting in front of every production request. Replay.io does time-travel debugging for browser sessions. HAR files hold requests and responses but nothing behind the server. VCR-style cassettes (Ruby’s VCR, Polly.js, go-vcr) record the outbound HTTP your code makes, which is one of the pieces you need. Keploy records API calls plus the dependency calls they make and replays them as tests with mocks, so it’s the closest relative. Sentry and similar error trackers attach rich context to a failure (stack, request data, breadcrumbs). You can read that context but you can’t run it.
The gap is a bundle that is server-side, scoped to one request, carries its data fixtures, and is written as an open format.
What Goes in the File
Middleware wraps each request in a recording context that notes every input the handler touched: the sanitized request, the resolved config and flag values, the git SHA and a hash of the lockfile, each database query with its result rows, each outbound HTTP call with its response, every clock reading, and the seeds behind random numbers and generated IDs. If the request ends in a 5xx or an uncaught exception, the context writes the bundle. Otherwise it drops the buffer. It’s tail-based sampling from distributed tracing, with payloads attached.
The file is a zip holding manifest.json, request.json, env.json, a fixtures/ folder of rows and a mocks/ folder of outbound calls. The manifest is the part that makes it a format:
{
"repro_version": 1,
"trigger": { "status": 500, "exception": "KeyError: 'tax_region'", "signature": "a91c7e" },
"code": { "git_sha": "3f9c2ab", "lockfile_sha256": "e4d1...", "image": "sha256:9b2e..." },
"clock": { "now": "2026-10-04T09:41:07.113Z", "frozen": true },
"random": { "seed": 8821107 },
"flags": { "new_tax_engine": true },
"db": { "engine": "postgres", "schema_sha256": "77ab...", "rows": { "customers": 1, "orders": 1, "tax_rules": 4 } },
"mocks": { "rates.example.com": 1 },
"unhooked": ["in-process lru cache"],
"redaction": { "rules": "redact.yaml", "self_check": "reproduced" }
}
Running it looks like this (a proposed CLI, so the names are illustrative):
$ repro run checkout-500.repro
checkout 3f9c2ab ... ok, lockfile hash matches
scratch postgres: schema 77ab ok, loaded 6 rows
mock server: 1 outbound host, clock frozen at 2026-10-04T09:41:07Z
replaying POST /v1/checkout
500 KeyError: 'tax_region' at pricing/tax.py:212
signature a91c7e matches the captured failure
REPRODUCED
The runner checks out the SHA (or pulls the image digest), loads the fixtures into a scratch database, serves the outbound mocks, freezes the clock, sends the request and compares the failure signature with the captured one. After a fix the same command should flip to a pass, and at that point the file has become a regression test for the suite built from traffic.
The fixtures go into a real scratch database, not a query cassette keyed on SQL text. A fix usually changes the queries, and a cassette fails on the first one it hasn’t seen. A database holding the rows the original request read can answer a new query as long as the rows are there, and a missing row tells you something: the fix reads data the failing request never touched. After the first reproduction the runner can also shrink the fixtures with delta debugging, dropping rows and replaying and keeping any smaller set that still fails. Six rows might become one, and that one, with its null tax_region, is usually the bug.
Reads and Redaction Are the Hard Parts
You don’t know in advance which request will fail, so the context has to record reads on every request. That means copying result rows into a buffer, which costs real memory on read-heavy endpoints. Cap bytes per request and mark a truncated bundle as incomplete rather than letting it pretend. Cap the process total, and ship a kill switch. Sampling can’t help, since you can’t sample the failure you haven’t had. The triggers matter too: 5xx and uncaught exceptions, plus a list of error codes, which is only precise if your errors carry stable codes (the error handling post covers how to build that). Write one bundle per error signature per hour, not one per occurrence; how you group errors decides which failures earn a file, and grouping by behavioral signature is the better way to decide.
Redaction has to keep referential integrity. Customer 4821 must become the same fake ID in the request, in every row and in every mock URL, so use deterministic pseudonyms from a keyed hash per bundle. The sharper problem is that the bug is often in the data. If the trigger was a name with an apostrophe or a trailing space, replacing it with user_17 hides the bug. Rules need a shape-preserving mode (letters to letters, digits to digits, punctuation, whitespace, non-ASCII characters and nulls left as they are), and the writer needs a self-check: replay the redacted bundle in a sandbox at capture time and confirm it still fails with the same signature. If it doesn’t, flag the bundle as “redaction changed the outcome” and leave the original on the server for someone with access.
Nondeterminism splits in two. Time, UUIDs and random numbers can be frozen or seeded if the code reads them through something you can wrap, which is easy in Python or Node and harder where the clock is a global call into the runtime. Concurrency races can’t be replayed from one request’s inputs, and v0.1 says so.
Environment drift comes next: OS, system libraries, locale, timezone. The lockfile hash says whether dependencies match, and an image digest is stronger because it pins the OS too. Proving that a described environment works from an empty machine is a discipline Preconfiguration applies to coding-agent setups: preconfig verify runs the setup on an empty Ubuntu container, then runs the project’s tests. A repro runner needs the same clean-room check.
Some state lives outside the database: an in-process cache, a rate limiter’s counter, a warm connection pool. A bundle can only hold what came through a hooked boundary, so hook the cache client and list the gaps in the manifest, as unhooked does above. Finally, copying production data to laptops breaks many companies’ data rules. Redact first, encrypt the file at rest with a key the team controls, put an expiry in the manifest that the runner enforces, and log who opens it. Treat even a redacted bundle as sensitive.
Version 0.1
Pick one framework. Python is the easier start, since the clock, the HTTP client and the database driver can all be wrapped at runtime. Capture Postgres or SQLite reads into fixtures, record outbound HTTP as mocks (the same idea as any mock server for integrations), freeze the clock, and read a redaction rules file. Refuse concurrency bugs, multi-service traces and binary protocols.
Write the manifest spec first: one page, a version field, and a rule that readers keep unknown keys. Adoption comes from the second writer and the second reader, not from the first tool. Agent systems are heading the same way, where a flight recorder in one portable file plays the role this plays for requests.
Steps to reproduce should be a command.