Fuzzing MCP Servers: Generate Bad Arguments From the Tool Schema and Watch What Breaks
You test an MCP server by chatting with it. Ask for the open bugs, get the open bugs, ship it. The model never sends limit: "ten", an empty path or a 200 KB query while you’re watching, so those paths stay dark until a real session hits one. Say it’s the limit. The handler throws, the framework wraps the exception in a 40 KB stack trace, and the model reads all of it, adjusts, and retries with the same bug in a new shape.
Every tool already publishes a grammar for its inputs: inputSchema, a JSON Schema. That’s what property-based testers feed on. Schemathesis does it for OpenAPI and GraphQL schemas, and RESTler, from Microsoft Research, fuzzes stateful REST APIs by working out which response feeds which later request. An MCP fuzzer applies the same idea to a new shape: discover the tools, generate arguments that obey the schema and arguments that break it, poke the protocol edges of the stateless spec, and record every finding as a seed you can replay. mcpfuzz below is an illustrative name for a tool you’d have to build.
The obvious product is a compatibility report, a grid of green and red cells. The better version treats the report as a by-product. The product is the shrunk, replayable failing input, because a finding nobody can reproduce gets closed. Expect competition. MCP Inspector, the official debugging tool, builds a form from each tool’s schema, which suits one call and not ten thousand (the general MCP guide has a debugging and testing section for the manual route). An OpenAPI fuzzer pointed at an MCP endpoint sees one JSON-RPC URL and nothing to enumerate. And anyone who has used Schemathesis could write a first version in a weekend. A newcomer stands out on protocol checks for the 2026-07-28 revision, safe defaults and shrinking. The read-only cousin of this tool is mapping an unfamiliar API by following its IDs.
What to Throw at a Server
Generation starts from inputSchema. For each property, take the boundary values for its type: minimum, maximum, one past each, zero, and an integer above 2^53, which a Node server rounds without telling anyone. For strings, send an empty one, one at maxLength and one just over, a 100 KB one, and the usual Unicode troublemakers (a null byte, a right-to-left override, a lone surrogate escape, an emoji outside the basic plane). For objects, drop each required property in turn, add unknown ones, and nest values a thousand levels deep to see whose parser gives up. Then break the types: a string where an integer belongs, null, an array for an object. Where you have real calls, mutate those instead; arguments from a recorded agent run beat anything generated from scratch.
Then check what comes back. A tool that declares an outputSchema owes the client a structuredContent that validates against it. A bad argument should produce an isError: true result or a JSON-RPC error, never a crash, a dropped connection or an HTTP 500, and the error text should help a model fix its next call: name the offending field and skip the stack trace. Good error bodies matter more when the reader is a model, because it acts on every word. Flag responses over a size threshold you set, since a tool that returns 2 MB to a model is a token bomb. On stdio, flag any stdout line that isn’t a JSON-RPC message; a stray debug print corrupts the stream.
The 2026-07-28 revision gives the protocol layer new things to break. Send a wrong MCP-Protocol-Version and expect a clean error rather than a crash. Drop it entirely and record what the server assumes. Send Mcp-Method and Mcp-Name headers that disagree with the body: a gateway may authorize on the header Mcp-Name: search while the server executes a body that says delete_repo, so a server should reject the mismatch. Send an unknown method and expect -32601, malformed JSON and expect -32700. Send a made-up Mcp-Session-Id, which the revision removed, and see whether the server ignores it. Test statelessness directly: run the same read-only calls in three shuffled orders against a fresh server each time and compare; a server that cares about order is hiding session state somewhere. Check cache hints too. If tools/list differs between two credentials but says cacheScope: public, a shared cache can serve one caller’s tool list to another. And when a server returns an input_required interim result, retry with the request state altered. The client echoes that state back, so it’s client-controlled input.
Last comes behavior under stress. Give each call a deadline and report hangs apart from slow calls. Fire the same read-only call in parallel and compare against the serial answer; differences mean shared mutable state. Throughput and latency percentiles belong to a benchmarking tool, which needs different statistics.
Security inputs belong in the corpus too. Tool arguments are often written by a model from text it read, a ticket or a web page, and some of that text was written by someone hostile. So ../../ goes in every path argument, shell metacharacters in anything command-shaped, SQL fragments in anything query-shaped. This is about testing your own server, and the problem is the oracle: a clean response doesn’t prove a traversal failed. Plant a canary instead, a file outside the directory the tool should serve, holding a unique string, and flag any result that contains it. Resource exhaustion also has a place in the OWASP API Security Top 10, and a tool that accepts a 50 MB string is an instance of it.
Brakes Come Before Cleverness
Fuzzing a delete_repo tool against production is a career event, so the default posture is refusal. The fuzzer won’t talk to a non-loopback address without an explicit --allow-remote, because the user, not the server, decides a target is disposable. Inside that boundary, tools marked destructiveHint: true are skipped, and so are tools carrying no annotations at all, because silence tells you nothing. Naming a tool with --tools overrides that for tools you vouch for. Annotations are untrusted hints, so they can only make the fuzzer more careful; a readOnlyHint: true doesn’t make a production server safe to fuzz. A --dry-run flag prints the tools and the case count per tool without sending anything. And every run opens by reporting what it skipped and why, so nobody concludes a server is clean when half of it was never touched.
Knowing What Counts as a Failure
The oracle problem is the real design work. A crash, a hang, an HTTP 5xx, a broken JSON-RPC envelope and an output that violates its own outputSchema are failures with no argument. Accepting an input the schema forbids is a warning at most, because JSON Schema allows extra properties unless the schema says otherwise, and a loose server is only buggy if the handler then misbehaves. An unhelpful error message is a warning too. Grade findings in tiers and fail the build on the first tier only, or people will switch the tool off.
Stateful tools make shallow fuzzing worthless. Fuzz update_issue with random IDs and every case bounces off a “not found” check before reaching the code that matters. RESTler’s answer, matching a producer’s output field to a consumer’s input, carries over: create_issue returns an id, update_issue takes an issue_id, and with an outputSchema the match is cheap to guess. Version 0.1 should settle for a hand-written setup block per tool and leave inference for later.
Flakiness is the other trap. Rerun every failure three times and report deterministic failures apart from intermittent ones. Derive each case from the run seed plus a case index, so 91a3:117 replays exactly. Then shrink, the way Hypothesis does: delete properties, halve strings, flatten nesting, push numbers toward zero, and keep a change only if the failure signature survives. The signature is the failure class plus the error text with digits and IDs stripped. Without it, shrinking wanders off to a different, duller bug.
$ mcpfuzz http://localhost:8808/mcp --seed 91a3
discovered 6 tools, fuzzing 2 (4 skipped: 3 destructive, 1 unannotated)
search_issues 412 cases 1 finding
get_file 380 cases 0 findings
FAIL search_issues uncaught exception returned as HTTP 500 (3 of 3 reruns)
seed 91a3, case 117, shrunk from 1.2 KB to 31 bytes
args: {"query":"bug","limit":-1}
error: "IndexError: list index out of range" (never mentions "limit")
replay: mcpfuzz replay 91a3:117
Scope for Version 0.1
The first version speaks HTTP and stdio, generates from inputSchema, runs the protocol checks for the 2026-07-28 revision, writes a Markdown or JSON report with seeds and shrunk reproducers, and has --skip-destructive on by default. It refuses AI-generated attacks, because a model inventing inputs adds back the nondeterminism you’re trying to remove from the test. It refuses load testing too. That’s another tool.
Fuzz your own server before a model does.