Turning Ten Minutes of Production Traffic Into an API Regression Suite
You’re about to refactor the billing endpoints. The service has a few dozen tests, mostly happy paths, and nobody trusts them to catch a changed rounding rule or a renamed field. Writing better ones by hand means reading every handler and inventing inputs. Meanwhile production receives thousands of real inputs a minute, and each one comes labelled with the response your current code gives.
Capturing that traffic is the easy part; a proxy, a packet tap or a log line with the body in it will do. The product is everything after. Ten minutes of traffic holds thousands of near-duplicate requests and a handful that exercise something different, and a tool is only useful if it can tell them apart. Cluster by endpoint, request shape and response shape. Pick representatives that cover the status codes and branches. Assert only on fields that are stable. Mock the downstream calls. What comes out is a few dozen characterization tests (Michael Feathers’ name for tests that pin down what code does today).
Characterization tests lock in current behavior, bugs included. That’s exactly what you want before a refactor and exactly what you don’t want as a specification, so the tool has to say which one it is. They also fill a different hole than the problems unit tests miss: they check what the service does over HTTP, with real shapes of data, not what one function returns for an input you imagined.
Who Already Does Part of This
Keploy is open source; it records API calls along with the dependency calls they make, then replays them as tests with mocks. Speedscale captures and replays traffic, and GoReplay shadows live HTTP traffic onto another environment. Twitter’s Diffy, from 2015, has the best trick in the field. It sends each request to a candidate build and to two copies of the stable build. Fields that differ between the two stable copies are noise (timestamps, generated IDs). Anything else that differs between stable and candidate is a regression. Nobody has to write an ignore list.
Judge every tool here with the same test: point it at ten minutes of real traffic and count the tests it hands you. If the answer is thousands, selection is still your problem. The gap is a small suite that lives in the repo as plain files, reads like something a person would review in a pull request, and arrives already denoised.
From Capture to Test
Normalizing comes first. Strip auth headers and cookies, drop volatile headers (Date, request IDs), sort query parameters, canonicalize JSON key order, and turn /orders/8812 into /orders/{id}. Then cluster. Key each request by method, route template, query parameter names, request body shape (key paths and types, no values), status code and response shape (key paths and types, with array lengths bucketed into empty, one and many), and count each cluster.
Selection is where the value is. Keep at least one representative per route and status pair. Prefer rare clusters over common ones, because a 422 seen twice says more about your code than the 10,000th 200. Then cap the total (say 40 tests). If you can run an instrumented build, add coverage guidance: replay the candidates, keep whichever one adds the most uncovered lines, repeat until nothing adds any.
Generation writes one plain-text file per cluster, in a real format. Hurl suits this because the output reads cleanly in a diff:
# cluster 7f3a: GET /orders/{id} -> 200, 3 items, coupon applied (412 in 10 min)
GET {{base}}/orders/{{order_id}}
Authorization: Bearer {{token}}
HTTP 200
[Asserts]
jsonpath "$.status" == "paid"
jsonpath "$.items" count == 3
jsonpath "$.total.currency" == "EUR"
jsonpath "$.coupon.code" exists
jsonpath "$.updated_at" matches /^\d{4}-\d{2}-\d{2}T/
# $.etag and $.items[*].reserved_until are volatile: not asserted
Noise filtering is the Diffy trick, adapted to a single build: replay each selected request twice against the same build, a moment apart, and diff the two responses. Paths that differ are volatile. Don’t delete them. Downgrade a volatile id from a value check to a shape check (“is a UUID”) so a field that vanishes still fails the test. Two replays a second apart share a blind spot, though: a field holding today’s date is identical both times and wrong tomorrow. Replay once more on a later day, or freeze the clock for the suite.
Mocks come from the same capture. Every outbound HTTP call a request made gets recorded next to its cluster and served back from a cassette, the pattern behind Ruby’s VCR, Polly.js and go-vcr. Match on method, host, path and normalized body, and fail loudly on an outbound call with no recorded match, since a changed outbound call is exactly the regression you want to see. The mocks and sandboxes post covers the mechanics of faking a dependency.
State and Noise Decide Whether Anyone Keeps It
A GET /orders/8812 only passes if order 8812 exists, and it exists because an earlier POST created it. Single requests can’t test that. Capture per session or per user, then generate chains: the POST response’s id goes into a variable (Hurl has [Captures] for this) and feeds the later path. A selector that picks individual requests will quietly pick GETs whose data doesn’t exist on the test machine, so the unit of selection has to be the chain. The same chains can be replayed under load, which is a different job (see benchmarking stateful workflows).
The database has to be in a known state too, and v0.1 doesn’t try to capture it. The suite assumes a seeded fixture or a restored snapshot that you provide. Say so in the generated README, because the alternative is a suite that passes on its author’s laptop and fails in CI. Recorded bodies also contain customers, so redact at capture time, and consistently: the same email must become the same fake email everywhere, or cross-references break. Packaging a single failed request goes deeper on redaction that keeps referential integrity.
Then there’s the rot. A captured 500 is not a spec, so skip 5xx responses by default and mark 4xx as expected. When a test fails after a change, a human has to answer “intended or regression?”, so give them an accept command that rewrites the expectations, the way jest -u rewrites snapshots. A test that fails twice with no code change loses its assertion on the offending path instead of staying red. And keep the suite small. A few dozen representatives that people read beat thousands of near-duplicates that people skip.
Ten minutes also has sampling bias. Tuesday at 2 p.m. misses the month-end job, the admin routes and the rare 404. So the tool should print what it didn’t see (“captured 23 of 41 documented routes, 6 of 19 error codes”) and treat the suite as a floor. Capture again at 3 a.m., or feed it the longer statistics a contract recorder collects, since both share one capture layer. The same record-and-replay shape works for an agent’s tool calls, as in replaying MCP traffic in CI.
Version 0.1
HTTP and JSON only. Capture through a proxy, cluster, select, and emit tests for one framework: Hurl if you want plain text, pytest if the team lives in Python. Filter noise with the replay-twice trick. It refuses gRPC, frontend sessions, and any database handling beyond mocked outbound calls.
Keep the suite small enough that a red test gets read.