Most MCP Token Waste Is in Tool Results: Put a Deterministic Reducer Between Server and Model
An agent calls a code-search tool and gets back a hundred hits. Each hit carries dozens of fields: node IDs, a URL for every related resource, avatar links, permission flags. The model needed three of them, a repo, a path and a snippet. The rest now sits in the context window for the remainder of the session, and the model reads past it on every later turn.
Tool definitions get most of the attention in agent token costs, and they’ve earned it. A server that exposes 150 tools puts 150 schemas in front of the model on every turn. Results are the other half of the bill, and they have fewer standard answers. An ordinary API client ignores the fields it doesn’t use, and ignoring is free. A model pays to read every token it’s handed.
The tool worth building is a deterministic reducer: a proxy between the MCP server and the model that applies per-tool rules to each result before the model sees it. Projections, row caps and binary stripping come first. Whatever is still too big goes to disk, and the model gets a short structural summary plus a way to ask for the rest. No model runs inside it, so the same result and the same rules always give the same output, and the savings can be measured instead of recalled from a good day.
Code Execution Covers Some of This
The definitions side has serious work behind it. Anthropic offers tool search, so the model finds tools on demand instead of carrying every schema. Cloudflare’s Code Mode turns MCP tools into a TypeScript API that the model writes code against. Anthropic’s write-up on code execution with MCP reaches results too: tools become code APIs loaded on demand, and intermediate results get processed in code instead of passing through the model. If you can run a sandbox and trust what the model writes into it, that’s the stronger design, because the reduction logic fits each task.
A reducer makes a narrower pitch. It works with the agent you already have, with servers you didn’t write and clients you can’t change, and its whole behavior sits in a file a reviewer can read. It’s the same idea as trimming API responses to the fields a client uses, applied to a reader that pays per token. Servers can trim their own output, but many are thin wrappers that hand back whatever the upstream API sent. Those are the ones you can’t fix at the source.
The cost is writing the rules, which pays off for the handful of tools an agent calls all day and not for the long tail.
How the Rules Would Work
The reducer can share a process with a small MCP proxy, since it needs the same spot on the wire: it sees every tools/call response and knows the server and tool from the request. Here’s a sketch, in illustrative syntax for a tool that doesn’t exist yet:
default:
max_bytes: 20k
always: [total_count, incomplete_results, next_cursor, truncated]
github:
search_code:
keep: "items[*].{repo: repository.full_name, path: path, url: html_url}"
max_rows: 25
filesystem:
read_file:
strip: [blob, image]
postgres:
query:
max_rows: 100
Projections are JMESPath expressions, which can reshape as well as filter. always lists the fields that survive any projection because they describe the part the model isn’t seeing: totals, cursors, truncation flags. strip swaps binary content for a one-line placeholder with its type and size.
Anything a rule cuts gets stashed. The original result goes to local disk under a hash of its canonical JSON, and the model gets the surviving rows with a note it can’t misread: “Showing 25 of 100 rows, fields repo, path and url. Full result stored as r_9f2c41; call reducer_expand with that ref and a row range or field list for more.” The result also carries a resource link for clients that support them, and reducer_expand covers the rest. The stash holds full results, so it’s keyed per caller and expires on a timer.
Every call writes one line to the accounting log:
{"ts":"2026-10-05T09:14:02Z","server":"github","tool":"search_code","bytes_in":412880,"bytes_out":3190,"tokens_in":103220,"tokens_out":798,"estimator":"bytes/4","cut":{"rows":75,"fields":["node_id","owner","..."]},"stash":"r_9f2c41"}
The numbers are made up. Bytes are exact. Tokens are an estimate, and models tokenize the same JSON differently, so a token savings figure is true for one tokenizer only. The log names its estimator and keeps the bytes beside it, so the figure can be recomputed for any model later.
An optional mode covers definitions with a few meta-tools (list categories, describe a tool, call a tool). It costs a round trip per new tool and duplicates tool search where the platform has it, so it ships off.
The cheapest reduction happens before the data becomes a result. A tool that returns rows makes the model do arithmetic over text. Precomputing keeps answers current in a SQLite file as data arrives, and an agent can ask for them over MCP instead of reading raw rows and doing the math itself. That fits predictable questions. A reducer fits a tool that’s a window onto somebody else’s API.
The Hard Part Is the Field You Dropped
A projection is a guess about what the model will ask next, and some guesses will be wrong. The expand path is what keeps a wrong guess cheap, which makes it more important than the projection syntax. If the model can’t get the dropped part back in one call, the reducer has traded tokens for wrong answers.
The worse failure is silent. Cut a hundred rows to 25 without saying so and the model reports 25 matches with full confidence. Every reduction has to say what it did, in the text the model reads, with counts. Pagination is the same trap: a projection that keeps items and drops next_cursor tells the model there’s nothing more to fetch.
Errors go through verbatim. A tool error is short and tells the model what to change, so trimming one deletes the instruction that would have fixed the next call. Results flagged isError: true pass untouched, along with any upstream notice about partial data.
Structured output makes this harder. A tool that declares an outputSchema returns structuredContent, and the spec suggests also serializing the same JSON into a text block for older clients. The reducer has to trim both copies identically. If the schema marks fields as required, a client that validates the reduced result will reject it, so the proxy has to rewrite the schema in tools/list to match the reduced shape. A proxy that does rename and hide is in that path anyway.
Determinism comes down to a few concrete rules. Keys keep a stable order. Rows are chosen by position, never sampled. Stash references come from a content hash, not a random ID or a timestamp. Anything looser breaks caching and replay, because the same call stops producing the same bytes.
Measurement is where reducers get dishonest. The README line everyone wants (“cut context by 73%”) usually comes from the one 8 MB response that makes the demo. Report the median per session on real workloads, with the 90th percentile beside it, and report the misses: how often the model called expand after a reduction, and how often it repeated a call with narrower arguments. A tool with a high expand rate has a bad rule. Then replay recorded sessions with and without the reducer and check that the tasks still finish. Savings that cost task success are a regression.
What v0.1 Does and Refuses
Version 0.1 sits on the wire for HTTP and stdio servers, applies keep, max_rows, max_bytes and strip rules per tool, stashes whatever it cuts, serves reducer_expand, passes errors through, writes the accounting log as JSON lines and offers a --dry-run that replays a recorded result through the rules and prints before and after. That last one is how you write rules without guessing.
It refuses to call a model. LLM summaries can come later as a labeled option, but a summary written by a model arrives in the one channel the main model trusts, so it can be wrong exactly where nobody checks. It also refuses to learn rules on its own, since a rule that drops fields should be reviewed by a person.
Once results are small, what stays in the window as a session grows is a database problem of its own. A reducer is also one of a family of small tools that shrink data near the source instead of storing and shipping all of it.
The model needs the answer. The payload can stay on disk.