MCP Tool Calls Are Plain HTTP Now: Route Them With a Small Proxy Before Buying a Gateway
A platform team runs four MCP servers for its coding agents: GitHub, internal docs search, a read-only Postgres server and the ticketing system. Someone wants per-tool rate limits and a log of which agent called what, so a gateway evaluation goes on the calendar. Nobody has asked the cheaper question yet, which is how much of that the reverse proxy they already run can handle.
Since July, quite a lot. The 2026-07-28 revision made MCP stateless at the protocol layer: the initialize handshake is gone, the Mcp-Session-Id header is gone, and every request carries its own protocol version, client info and capabilities. Two new headers, Mcp-Method and Mcp-Name, copy the JSON-RPC method and (for a tool call) the tool name out of the body and into headers, where proxies already look. Without sessions, an MCP server behaves much like any other HTTP API, which is why the difference between an API and MCP now comes down to who reads the docs. Before the revision, a proxy needed a body-parsing module to learn which tool was being called, and usually session affinity to keep each client on the backend holding its session. Now stock nginx covers the basics:
log_format mcp '$time_iso8601 $remote_addr $http_mcp_method '
'$http_mcp_name $status $request_time $bytes_sent';
limit_req_zone $http_mcp_name zone=per_tool:10m rate=20r/s;
upstream github_mcp { server 10.0.1.10:8080; server 10.0.1.11:8080; }
server {
listen 8080;
access_log /var/log/nginx/mcp.log mcp;
location = /github/mcp {
limit_req zone=per_tool burst=40;
proxy_buffering off; # responses can stream back as SSE
proxy_pass http://github_mcp/mcp;
}
}
Those few lines give you a per-tool rate limit and a metering log keyed on tool name. The upstream block round-robins across two backends with no stickiness, which would have broken sessions before July. If your agent clients let you configure one server entry per backend, per-server paths like this cover a fair share of what a gateway demo will show you.
What a Stock Proxy Can’t Do
The gaps are specific, and every one of them involves reading MCP payloads. nginx can route a tools/call by header, but it can’t answer tools/list or server/discover for one endpoint that fronts four servers, since that means calling each upstream and merging the results. It can’t rename or hide a tool, which takes rewriting JSON in both directions. It trusts Mcp-Name without checking the body. It can’t cap a response without cutting JSON mid-string, it can’t launch a local server that only speaks stdio, and it has never heard of an input_required result.
That residue is what a small MCP-aware proxy should cover. The catch is that the gateway category is already crowded. Docker ships an open-source MCP Gateway in its MCP Toolkit; IBM has ContextForge; Microsoft has an MCP gateway for Kubernetes; agentgateway, a Rust proxy for agent traffic, is now a Linux Foundation project. Kong’s and Envoy’s AI gateways support MCP, there’s MetaMCP, and small bridges such as supergateway and mcp-proxy put stdio servers on HTTP. Across the category you’ll find aggregation, transport bridging, auth, logging, rate limits, caching and observability, and some projects advertise loading tool definitions lazily to save context.
Size and speed won’t separate a newcomer from that field; a benchmark lead is easy to lose and hard to sell. What’s still rare is a config you can reason about. So the proxy worth building is nginx-shaped: one binary, one file you can read top to bottom, and a few MCP-aware features chosen with care. It should also answer two questions about its own config: which upstream handles tools/call for code_search and why, and what a pending change would reroute before anyone applies it.
There’s a working model for that in plain HTTP. BareProxy, a small Go web server and reverse proxy, keeps regular expressions and scripting out of its routing so it can explain itself: bareproxy explain says which rule matched a URL and why, and bareproxy plan lists which existing requests would change hands under a new config. It’s an HTTP proxy and knows nothing about MCP. The design carries over anyway. A config for the MCP version might read like this (illustrative syntax for a tool that doesn’t exist yet):
listen 127.0.0.1:9000;
access_log /var/log/mcp/access.jsonl;
max_result 512k;
upstream github { http https://mcp-github.internal/mcp; }
upstream docs { stdio /usr/local/bin/docs-mcp --index /srv/docs; }
expose github {
tool search_code as code_search { limit 30/min per_user; }
tool get_file as read_file { max_result 128k; }
hide delete_repository update_branch_protection;
}
expose docs {
tool * prefix docs_;
}
$ proxy explain tools/call code_search
code_search -> upstream github, tool search_code
rule expose github: tool search_code as code_search (proxy.conf:9)
limit 30/min per user, keyed on the token subject
cap 512k, inherited from max_result (proxy.conf:3)
$ proxy plan proxy.conf.next --since 7d
rename docs_search -> docs_lookup 2,140 logged calls used the old name
route read_file: github -> github-ro 8,412 logged calls change upstream
cap read_file 128k -> 64k 31 logged calls would have gone over
plan reads the proxy’s own access log, so a rename arrives with a count of the calls still using the old name. The same nginx-style shape works one layer down, where an LLM router driven by one config file sorts model requests by token count, privacy tag and cost ceiling.
Where It Gets Hard
Renames and hides have to hold everywhere at once. A renamed tool must show its new name in tools/list, get translated back in the tools/call body before it goes upstream, and match the Mcp-Name header the client sends. Two upstreams that both expose search make renaming mandatory. A hidden tool has to be refused on call as well as dropped from the list, because a client can call any name it likes. Merged lists carry a quieter trap. Once the list depends on who’s asking, a result the upstream marked cacheScope: "public" isn’t public anymore and has to be rewritten as private. A merged list also stays fresh only as long as its shortest-lived part, so its ttlMs is the smallest of the upstream values.
Header routing needs body validation. Picture a request with Mcp-Name: search_code in the header and "name": "delete_repository" in the body. A proxy that checks policy against the header and forwards the body has just let the upstream run a blocked tool; it’s the JSON-RPC version of HTTP request smuggling. So the proxy parses every body anyway and rejects any request whose header and body disagree. Headers make routing cheap. Policy still has to read the body.
Multi Round-Trip Requests must pass through intact. When a tool needs more input, the server returns an interim result with resultType: "input_required" plus state the client echoes back on the retried request. The proxy mustn’t treat that interim result as final, cache it or trim the state. It also has to decide whether the retry counts against the rate limit, and counting it twice punishes exactly the tools that stop to ask.
Auth is where scope creep starts. The new revision replaces Dynamic Client Registration, now deprecated, with Client ID Metadata Documents, and per-user limits need a verified identity. Forwarding the agent’s token to every upstream is the easy answer and the wrong one, since a token minted for one server shouldn’t open another. (This MCP guide is a reasonable map of the transport and security pieces.) The v0.1 answer is narrow: verify incoming tokens against one issuer to get an identity, and hold separate credentials for each upstream.
Stdio servers need a bridge, and bridges have sharp edges. Servers written against earlier revisions still expect the initialize handshake, so the bridge performs it for them. Sending several clients through one child process means remapping JSON-RPC ids so they can’t collide. And a stdio GitHub server launched with one engineer’s token in its environment is that engineer’s server: put it behind a team proxy and everyone quietly inherits that engineer’s access.
Size caps can’t cut bytes. A JSON result truncated at 512 KB is something no client can parse. In v0.1, an oversized result becomes a tool error (isError: true) that states the size and the cap, so the model can narrow its request. Trimming results field by field is a different job, and it belongs in a deterministic reducer between server and model.
Staying small is the hardest part, because every user wants one more feature. The rule that keeps explain honest is that routing stays declarative, with no scripts and no regular expressions. Extensions run only after a route is chosen, and the config declares them, so explain can name them in its output. Caching is the obvious first extension, and a caching proxy for tool calls has enough hard problems of its own to stay an opt-in module.
What v0.1 Ships
The first release does HTTP and stdio upstreams, routing on Mcp-Method and Mcp-Name with body validation, a merged tools/list, rename and hide, per-tool and per-user rate limits, a response size cap and a structured access log in JSON lines. explain ships on day one because it’s the reason the project exists; plan follows as soon as there’s an access log to replay. It refuses dashboards, a database, a server registry and orchestration of any kind. If a feature needs state that outlives a request, it lives in a separate process.
The advice from the build-versus-buy gateway debate applies here too: gateway adoption should lag complexity. Four MCP servers and a handful of agents aren’t there yet.
Buy the gateway when the config stops fitting on one screen.