Infer Your API's Real Contract From Traffic, Then Diff It Against the Docs
Your OpenAPI file marks shipping_address as required on GET /orders/{id}. After Tuesday’s deploy, about one response in 300 leaves it out, because a new code path for orders created by the import job skips the field. The docs still say required. The tests still pass, since nobody wrote one for import-job orders. A client with a strict deserializer starts throwing on 0.3% of order pages, and the first report you get is “sometimes the order screen is blank”.
The documented schema is a claim. Traffic is evidence. A recorder that sits on the response path and keeps per-field statistics (how often a field is present, how often it’s null, which types and formats it takes, which values it takes and how many times each) gives you the contract you actually serve. The most useful thing it produces is the diff between that contract and your OpenAPI document: “documented required, missing in 0.3% of responses since Tuesday”, or “documented enum of four values, six observed”.
That diff matters more than the docs do, because clients get written against what an API returns. Hyrum’s Law says that with enough users, every observable behavior gets depended on by somebody, which is how every accident in your API becomes a contract. A spec diff in CI catches the breaking changes you meant to make. This catches the ones you didn’t.
What Exists Already
Akita Software built a product that inferred API specs by watching traffic; Postman acquired it in 2023. Optic is an open-source tool for diffing and linting OpenAPI files, and it could also update a spec from captured traffic. mitmproxy2swagger turns captured mitmproxy flows into an OpenAPI document. genson and quicktype infer JSON Schema or types from sample documents. Pact comes at the problem from the other end: consumers declare what they depend on, and the provider verifies it.
Most of these end in a spec, one document holding the union of everything seen. A union can’t say that a field shows up in 99.7% of responses; “optional” is the only word a schema has for it. Whether 0.3% is a quirk or a regression depends on the number, on when it started, and on which clients get the short version. Doing this once by hand is a decent afternoon (the field completeness audit is that afternoon). The gap is a recorder that keeps the counts, compares them over time, and is cheap enough to leave on every endpoint.
The Recorder Keeps Counts, Not a Union
Capture can happen in a reverse proxy, in framework middleware, in a sidecar, or from a HAR file for a one-off look. Whatever you pick needs response bodies, a size cap, and a hand-off that keeps inference off the request path so it never adds latency.
Statistics are keyed by method, route template, status code and API version (from the path or a header). A 404 and a 200 have different contracts, and /v1/orders and /v2/orders shouldn’t blur together. Route templates come from the OpenAPI paths when you have them; otherwise segments that look like integers or UUIDs collapse to {id}. Inside a body, array elements collapse to [], so $.items[].sku is one field, and its presence is measured against the number of items, not the number of responses.
For each field the recorder keeps a handful of counters: times seen, times its parent was seen, nulls, one count per JSON type, one per detected string format (uuid, date-time, email, URL), numeric and length ranges, and distinct values with counts up to a cap of about 50. Past the cap it flags the field as high cardinality and stops tracking values. Here’s one field from a hypothetical run (a sketch of the data, not output from a shipping tool):
endpoint: GET /orders/{id}
status: 200
field: $.shipping_address
parent_seen: 48211
seen: 48067
null: 0
types: { object: 48067 }
docs: { required: true }
first_missing: "Tue 14:02Z"
“Required” gets inferred with a threshold and a floor. A field counts as always present only after it has appeared in every response across a few hundred samples. An enum candidate is a field whose distinct count has stopped growing while the sample count keeps climbing: a status field that shows the same six values across 50,000 responses is an enum, while a field that showed 40 distinct values in the first 100 responses and 41 in the next 50,000 is an identifier.
Drift rules fall out of the counters: a new field, a vanished field, a presence rate that moved, a new enum value, a type that changed (a number became a string), a format that changed. Compared against the OpenAPI file, the same counters give the findings that matter most. This is the CLI the post is proposing, with made-up numbers:
$ contract diff --spec openapi.yaml --window 7d
GET /orders/{id} 200
DRIFT $.shipping_address documented required; missing in 144 of 48,211 responses (0.30%), first miss Tue 14:02Z
DRIFT $.status documented enum of 4; observed 6 (refunded: 212, partially_shipped: 31)
INFO $.customer.age documented integer; observed integer or null (null in 8.4%)
INFO $.gift_note undocumented; present in 2.1%
Run the comparison backwards and you get documented endpoints that never see a request. That’s its own problem, and traffic alone only answers half of it. The capture layer is also the one you’d use to generate regression tests, so building one gets you most of the other. With no traffic yet, active exploration can seed a first picture.
Telling Optional From Broken
Privacy comes first. A recorder that stores payloads is a second copy of your production data with weaker access controls. Store counters and salted hashes, never raw bodies. Keep example values only for fields that qualify as enums, cap how many you keep, and run format detection to redact obvious personal data (emails, phone numbers, tokens) even there. Statistics can leak too: a distinct-value list on a name field is a customer list.
Polymorphism wrecks naive presence rates. If a payment response carries card details in 60% of responses and bank_transfer details in 40%, no field is optional; the discriminator decides. Before calling a field optional, check whether a low-cardinality sibling explains its presence, and report it as “present when type is card”. Heterogeneous arrays need the same treatment: cluster items by shape before counting.
Low-traffic endpoints never reach significance. After n clean samples, the 95% upper bound on a miss rate is about 3/n, so a thousand clean responses only tell you the rate is under 0.3%. Below the sample floor, the honest report says “not enough data”. Time windows add a trap of their own. A statements endpoint might add closing_balance only on the last day of the month, so a week of traffic calls it absent. Keep daily buckets, treat the first cycle as a provisional baseline, and don’t alert on seasonal fields until you’ve watched one full period.
Then comes the judgment call: is 0.3% an optional field or a bug? The tool can’t know, but it can bring evidence. An optional field has a stable rate; a bug shows a step change, usually lined up with a deploy. Test the recent window against the baseline with a binomial test plus a minimum absolute change, so a drift from 99.9% to 99.8% on a million responses doesn’t page anyone. Attach what correlates with the misses (clients, sibling values, release) and let a human pick between “document it as optional” and “fix it”.
Version 0.1
A small reverse proxy, JSON responses only, per-endpoint statistics in SQLite, and a diff command that reads an OpenAPI file. Add a CI step that runs the diff against a committed baseline and fails only on drift nobody has acknowledged, so the check doesn’t go red on day one over old sins.
It refuses request bodies at first. Traffic shows what clients send, which differs from what the server would accept, so a request contract inferred from it describes your clients rather than your API, and request bodies carry the most sensitive data anyway. It also refuses GraphQL (introspection already hands you the schema) and binary formats.
Run it on one endpoint for a week. You’ll find something.