The New MCP Spec Caches Tool Lists but Not Tool Calls, and a Caching Proxy Fills the Gap
An agent works through a ticket and asks a docs-search tool the same question on turn 3, again on turn 14, and again the next morning in a fresh session. Each repeat costs a charge against the upstream’s rate limit and a wait with the model idle, for an answer that never changed. Agents retry after errors and start every session by looking up what the last one already knew.
The 2026-07-28 revision of MCP added a caching vocabulary, and it stops one step short of this case. Six results now carry ttlMs and cacheScope: tools/list, prompts/list, resources/list, resources/templates/list, resources/read and server/discover. The fields work like HTTP’s Cache-Control. ttlMs is a freshness hint in the spirit of max-age, where 0 means stale at once, and cacheScope is public or private. Interim input_required results can’t be cached. Results from tools/call aren’t on the list at all, and it’s easy to see why: a protocol can’t know whether repeating a call is safe. Only the operator does. (Some SDKs also default ttlMs to 0, which makes a compliant client refetch the tool list on every turn.)
That gap is a job for a caching proxy. It sits beside a small MCP proxy or inside one, since it needs the same view of every call, and it takes a policy from the operator that says which calls may be served again. The version worth building denies by default, keys on identity and tells the model how old an answer is.
Prior Art Gets You Halfway
Since the spec went stateless, MCP over HTTP behaves much like any other API, so HTTP’s caching rules are the right thing to borrow. RFC 9111 defines freshness and validation, and RFC 5861 adds stale-while-revalidate and stale-if-error. LiteLLM, Portkey, Helicone and Cloudflare AI Gateway all cache responses, but they cache model calls, which is a different problem with its own payoff pattern. For plain HTTP, BareProxy’s Cache module shows the opt-in shape: it keeps responses that say they can be cached and purges by path prefix. MCP needs the same idea with an operator’s policy standing in for the response headers.
A TTL on a tool name is the easy half of caching. The other half decides whether the cache is safe.
A Policy File and a Key
Everything is off until a rule turns it on. A rule names a server and a tool, and the cache key is built from the server’s identity, the tool name, the canonicalized arguments and, for private results, the caller’s identity. Here’s a sketch of the policy, in illustrative syntax for a tool that doesn’t exist:
[defaults]
mode = "deny" # nothing is cached unless a rule says so
age_note = true # tell the model when it's reading a cached answer
[[rule]]
server = "docs"
tool = "search_docs"
ttl = "6h"
swr = "1d" # serve stale while refreshing in the background
scope = "public"
ignore = ["trace_id"] # arguments that don't change the answer
skip_if = { version = "latest" }
[[rule]]
server = "github"
tool = "list_issues"
ttl = "2m"
scope = "private" # the key includes the caller's identity
purge_on = [
{ tool = "create_issue", match = ["owner", "repo"] },
{ tool = "add_comment", match = ["owner", "repo"] },
]
Canonicalization means sorted keys, no insignificant whitespace, and defaults filled in from the tool’s inputSchema, so an omitted argument and an explicit default hit the same entry. It stops there. Treating “react hooks” and “hooks in react” as one query is semantic caching, and a near-miss that returns a different answer is worse than a miss, so v0.1 matches exactly. The proxy reads the method and tool name from Mcp-Method and Mcp-Name, but the arguments live in the body, so it parses the body and checks that the two agree before any lookup. Otherwise a request could ask for one tool and be served another tool’s cached answer.
The annotations readOnlyHint and idempotentHint are hints from the server, and clients shouldn’t trust them from servers they don’t control. The proxy uses them only to prompt suggestions. A person’s rule is the only permission.
A suggest command goes further by watching traffic. It logs each call’s key and a hash of the result, and for every tool it measures how long identical calls kept returning identical results:
$ mcpcache suggest --since 7d
github.get_repository 412 of 1030 calls repeated changed after 3h40m suggest ttl 30m
docs.search_docs 2206 of 3115 calls repeated never changed suggest ttl 6h
web.fetch_url 180 of 940 calls repeated answers differ no suggestion
The numbers are made up. The suggested TTL sits well under the shortest stable span seen, and a person approves it before it enters the policy. A tool whose repeats come back different gets no suggestion.
The six cacheable results need handling too. The proxy honors the server’s ttlMs and cacheScope unless the operator overrides them. A server whose SDK defaults ttlMs to 0 can be told its tool list is good for ten minutes, and the proxy rewrites the hint on the way out so the client’s own cache starts working. It never widens a scope: a private result stays private, and a merged list that depends on the caller never leaves as public.
The two RFC 5861 extensions earn their keep here. swr covers slow tools by answering from the stale entry at once and refreshing behind it. stale-if-error gives a rough offline mode, where a dead upstream means an old answer, flagged as old, for a bounded window instead of a failed tool call. Storage is the dull part, and an embedded key-value file or SQLite does it. Keeping responses in SQLite already gets you a cache, an offline mode and a history from one table.
Where the Cache Goes Wrong
A shared cache is a data leak waiting for a lazy key. A list_issues result depends on who’s asking, so a key without the caller’s identity serves one user’s private issues to another. Private is the default, the identity comes from the verified token and never from an argument the client sends, and public needs an explicit rule. A server labeling its own result public isn’t permission.
Time-dependent tools defeat TTLs. “Latest”, “today”, “now” and search results over a corpus that’s still changing all return answers whose useful life is shorter than any default. That’s why the policy refuses by argument value as well as by tool name: version = "latest" skips the cache, version = "4.2" doesn’t.
Read tools have side effects. A call may bump a view counter, mark a message read, spend rate limit or write an audit record, and a cache hit skips all of it. Sometimes that’s the point, since the rate limit is why you’re caching. Sometimes it breaks something, and readOnlyHint can’t tell you which. Someone reads the tool’s docs once per rule.
Multi Round-Trip Requests complicate keys. A call that returns input_required and then completes after the client echoes state back produces a final result that depends on that state, including whatever the user answered. Interim results aren’t cacheable, and the simplest safe rule is that no call carrying request state gets cached, interim or final.
Poisoning is the quiet risk. A compromised or buggy upstream that returns a bad answer with a long TTL serves it to everyone until the entry expires, and a prompt injection inside a cached result keeps working as long as the entry lives. Cap TTLs per server, never cache isError: true results (a transient failure shouldn’t turn into a lasting one), skip results over a size cap, and keep a purge command by server, tool or key. Log the result hash with every hit so a bad entry can be found afterward.
The model needs to know too. Read a 40-minute-old answer as live and it reasons from stale facts with full confidence. Put the age in _meta for clients that read it and, when age_note is on, as a first line of text, because text is the only part sure to reach the model. That line changes the bytes of the result, so anything comparing cached output with upstream output has to strip it first. Freshness is a property the model should be able to see, which is the premise of a database where facts expire.
A hit rate says how often the cache answered, not whether it answered correctly. For that, revalidate a small sample of hits in the background and count the ones where the upstream now disagrees. A tool with a high mismatch rate has a TTL that’s too long.
What v0.1 Does and Refuses
Version 0.1 is an exact-match cache with default deny. It does per-tool TTLs, private scope keyed on verified identity, purge-on-write rules, stale-while-revalidate and stale-if-error, cache age in _meta with an optional text line, correct handling of ttlMs and cacheScope on the six cacheable results with operator overrides, storage in a single file, a hit and miss log with result hashes, and suggest.
It refuses semantic matching. It refuses to cache errors or anything that carried request state. It runs on one node and won’t apply a learned TTL until a person approves it.
Cache only what you’ve named.