Benchmarking MCP Servers and Stateful API Workflows: Latency, Throughput and Tokens per Call
Point wrk at an MCP server and you get a clean report: every response a 200. Some of those 200s carry isError: true results, and nothing in the output says what the tool definitions cost the model on every turn. The tool is measuring requests per second. An agent runs a workflow, and what the server costs it is latency per call times calls per task, plus the tokens each response and each definition puts into its context.
Now that MCP is plain stateless HTTP, any load generator can hit an endpoint. They just measure the wrong things. The tool worth building reports per-tool latency percentiles, the cost of server/discover and tools/list, error rates split by JSON-RPC code and by isError results, and payload sizes with their token cost. Tool definitions are paid on every turn, so a bloated tool list is a recurring context tax that no latency number captures. Give the same engine variables captured between steps and it covers stateful API workflows too: authenticate, create an object, query it, update it, delete it. mcpbench below is an illustrative name for a tool you’d have to build.
The competition is real. wrk, hey, oha and vegeta are good at hammering one endpoint, and wrk2 added a constant-rate mode that corrects for coordinated omission (more on that below). k6 runs JavaScript scenarios and Artillery runs YAML scenarios with captured variables, and both can model arrival rates, so the stateful half of this idea is already served by mature tools. Hurl chains requests with captures and asserts in plain text but isn’t a load tool. A small newcomer won’t out-feature k6. It can win on zero-config MCP awareness (discovery, a per-tool breakdown, the right definition of an error) and a scenario file short enough to read on one screen.
What an MCP-Aware Run Reports
A run starts the way a client would. It calls server/discover, then tools/list, timing each and recording the bytes. Those calls get their own lines because a client pays for them again whenever the server’s cache hints (ttlMs) run out. A stdio server adds a startup cost, spawn to first answer, which the stateless revision makes easy to measure because there’s no handshake to wait through. Then the run calls the chosen tool at a fixed arrival rate, taking arguments from a file. Identical arguments every time would measure the server’s cache, not its work. Each request carries Mcp-Method: tools/call and Mcp-Name: search, so anything routing on those headers sees realistic traffic.
A run would print something like this (the numbers are invented to show the shape of the report):
$ mcpbench https://staging.example.com/mcp --tool search --args-file queries.jsonl --rate 50 --duration 60s --warmup 10s
server/discover 1 call 38 ms 2.1 KB
tools/list 5 calls p50 12 ms 31.4 KB ~7.9k tokens (est.)
tools/call search 3000 calls p50 41 ms p90 88 ms p99 310 ms max 1.2 s
errors 9 (0.3%): 8 isError results, 1 JSON-RPC -32603
result size p50 6.2 KB, max 71 KB ~1.6k tokens per call (est.)
client lag p99 0.4 ms latency measured from scheduled send time
The size lines matter more than they look. That 31 KB tool list sits in the model’s context on every turn of every conversation that uses the server, so a long task pays for it again and again. Once the bench has shown you the number, a reducer between server and model is the usual fix. Tokens are the soft number, though. A tokenizer belongs to a model, so the tool reports bytes as the measurement and tokens as a labeled estimate with a configurable ratio, or a real tokenizer if you plug one in.
Measuring Without Lying to Yourself
Start with the load model. A closed model runs N workers, each sending its next request when the last one returns. That’s what --concurrency 100 gives you, and it flatters slow servers: when the server stalls, the workers stall with it, and the offered load drops exactly when it should rise. An open model sends requests at a fixed arrival rate whatever the responses do, the way independent users show up. Real traffic is open, so the default should be a rate, with fixed concurrency as the option. Agents add a twist. Each agent is a closed loop of one, since it waits for a tool result before its next call, but the fleet of agents arrives openly. So the natural unit is a scenario instance, started at an arrival rate and running its own steps in sequence.
Coordinated omission is the flaw closed-loop tools bake in, described by Gil Tene. If a request stalls for ten seconds, a closed-loop generator sends nothing during those ten seconds, so it never records the requests that would have been sent and delayed. The tail of the distribution vanishes from the data. The fix wrk2 applies is to schedule requests on a fixed timeline and measure each latency from the moment the request was supposed to start. The bench does the same.
Throw away a warmup window. Connections, TLS sessions and caches are cold at the start, and those samples describe the first ten seconds of the run, not the server. Report cold-start cost on its own line instead of burying it. Then report percentiles, never averages, kept in an HdrHistogram so memory stays fixed and the tail keeps its resolution. Print the sample count beside every percentile, because a p99.9 from a few thousand samples is a handful of requests.
Measure your own lag too. The report should say how late each send was against the schedule; if the client can’t keep up with the rate, you’re benchmarking the client, and the output should say so loudly instead of printing flattering numbers. Run it from a machine that isn’t also running the server. And never aim it at someone else’s server, or your own production, without permission. Fifty requests a second can knock over a small service, and a scenario that creates objects pollutes real data.
A benchmark says how slow, not why. For why, profile before optimizing, because most of the time usually goes to the database and the network. When one percentile looks wrong, splitting a single request into DNS, connect, TLS, server wait and transfer shows where it went, and server-side logs, metrics and traces show why. If the server sends Server-Timing headers, the bench can record those phases as well.
A Scenario File Short Enough to Read
The second mode reads a YAML scenario. Each step is one request, capture pulls values out of the response for later steps, and always marks cleanup that must run even when an earlier step failed. Variables come from three places: $env for secrets, captures, and the engine’s own run and n (the run ID and the instance counter).
name: ticket-lifecycle
base_url: https://staging.example.com/api
arrival_rate: 20/s # new scenario instances per second
duration: 2m
warmup: 15s
steps:
- {name: login, post: /auth/token, json: {key: $env.API_KEY}, capture: {token: $.access_token}}
- {name: create, post: /tickets, bearer: "${token}", json: {title: "bench ${run}-${n}"}, expect: 201, capture: {id: $.id}}
- {name: read, get: "/tickets/${id}", bearer: "${token}"}
- {name: update, patch: "/tickets/${id}", bearer: "${token}", json: {status: closed}}
- {name: delete, delete: "/tickets/${id}", bearer: "${token}", always: true}
The report comes per step and for the whole instance. Four things make this harder than it looks. Cleanup is the first. A run that creates thousands of objects and dies halfway leaves them behind, so every object carries the run ID in its title and a cleanup subcommand sweeps by it. State growth is the second: a list endpoint slows down as the benchmark fills the table, so a run that deletes as it goes, like the scenario above, and one that doesn’t will report different numbers for the same server.
Rate limits distort everything. A 429 is the limiter’s decision, not a slow response, so count it as its own class and don’t retry automatically, because retries change the arrival process. Ask for the limit to be lifted for the benchmark key; otherwise you’re measuring the rate limiter. Finally, Multi Round-Trip Requests and slow tools break the one-request-one-latency assumption. A call that returns input_required is two round trips, so the scenario needs a scripted answer for the interim input, and the report shows the logical call with its round trips and each leg. A tool that runs for minutes by design needs its own deadline and a time-to-first-progress number, since a p99 of minutes is the tool working as intended.
Numbers from different machines are unrelated to each other. Print the client’s CPU count, OS and tool version in every report, seed the argument picker, and compare a candidate against a baseline from the same machine, back to back.
Scope for Version 0.1
The first version has an MCP mode (discovery plus one tool-call pattern), an HTTP scenario mode with captures, an open-model arrival rate, HdrHistogram percentiles and JSON output. It refuses distributed load generation, because one machine covers the common case of checking a server before you ship it and the multi-machine version is a different product. It refuses a UI. It refuses an embedded scripting language too: when a scenario needs a loop, that’s the day to reach for k6. A fuzzer shares the discovery code but not the job, so fuzzing MCP servers stays a separate tool.
Measure the task, not the endpoint.