Below you will find pages that utilize the taxonomy term “Testing”
Package a Failed API Request Into One File Anyone Can Replay Locally
A customer’s checkout returns a 500. Support pastes the request ID into the ticket, and the engineer on call finds the log line: KeyError: 'tax_region' in the pricing module. They send the same request locally and get a 200. Of course they do. Their database has no customer with a null tax region, the feature flag that routes to the new tax engine is off in development, the rates service answers differently today, and the clock is a day later. The bug is a function of all of that, and the ticket contains none of it.
Record an Agent's MCP Traffic Once, Then Replay It in CI Without Servers or Credentials
Your agent test passes on your laptop and fails in CI. The laptop has a token for the issue tracker’s MCP server; CI doesn’t, and shouldn’t. Even with a token, the tracker listed three open bugs yesterday and lists four today, so the agent’s summary changes and the assertion on it breaks.
Web developers dealt with this years ago. VCR (Ruby), Polly.js, Betamax and go-vcr record real HTTP responses into a file the first time a test runs, then serve them from that file on every run after. The file is called a cassette. MCP needs the same thing: a proxy that sits between an agent and its MCP servers, writes every exchange into one cassette, and plays it back later with no server, no credentials and no network. mcprec below is an illustrative name for a tool you’d have to build.
Turning Ten Minutes of Production Traffic Into an API Regression Suite
You’re about to refactor the billing endpoints. The service has a few dozen tests, mostly happy paths, and nobody trusts them to catch a changed rounding rule or a renamed field. Writing better ones by hand means reading every handler and inventing inputs. Meanwhile production receives thousands of real inputs a minute, and each one comes labelled with the response your current code gives.
Capturing that traffic is the easy part; a proxy, a packet tap or a log line with the body in it will do. The product is everything after. Ten minutes of traffic holds thousands of near-duplicate requests and a handful that exercise something different, and a tool is only useful if it can tell them apart. Cluster by endpoint, request shape and response shape. Pick representatives that cover the status codes and branches. Assert only on fields that are stable. Mock the downstream calls. What comes out is a few dozen characterization tests (Michael Feathers’ name for tests that pin down what code does today).