LLM Response Caching Pays Off in CI, Evals and Agent Retries, Not in Chat
A repository has 300 integration tests that each send a prompt to a model. The suite runs on every push to every branch, then again in the merge queue. Most of those runs don’t touch a prompt (the diff was a stylesheet or a migration), so the model receives the same 300 requests it saw an hour earlier, writes roughly the same 300 answers, and the provider bills every one. Nobody chose that. It’s what happens when a test calls a live API.
A response cache fixes this case, and the case is narrower than the usual pitch. “Put a proxy in front of the model and cut your bill” sounds like it’s about production chat, because that’s where the traffic is. Chat is where it does least. An exact-match cache hits only when a request repeats byte for byte, and live conversations almost never do: every turn carries the full history, and the first message varies with whoever typed it. The requests that repeat come from machines. Test suites, eval runs, an agent restarting after a crash, a queue that delivers the same job twice, a developer running one script forty times an hour while fixing the parser behind it. The tool worth building is a boring exact-match cache aimed at those paths, with a key you can read and a report that says what it saved.
Two Things Called Caching
Provider prompt caching and response caching share a word and little else. Anthropic’s version is explicit: you place cache breakpoints in the request with cache_control, and later requests that share the prefix up to a breakpoint read it from cache, billed at a reduced rate. OpenAI applies prompt caching automatically once a prompt reaches 1,024 tokens. Both cache the prompt prefix. The model still runs, so you pay for every output token and for whatever follows the cached prefix, and generation takes as long as it takes.
A response cache skips the call. A hit is a disk read that bills no tokens and returns in milliseconds. The two stack well, since a request that misses the response cache can still hit the provider’s prefix cache, which means the proxy must never disturb prefix stability by reordering tool definitions or “cleaning up” a system prompt.
None of this is new. GPTCache, from Zilliz, is a library built around semantic caching of LLM calls. LiteLLM, Portkey, Helicone and Cloudflare AI Gateway all offer response caching among many other features; the gateway section of this review of AI platforms for API work covers what else they do. If you already run one of them, switch its cache on before writing code. The opening for a separate tool is a single binary you can drop into a CI job or an eval harness without adopting a gateway, with key rules strict enough to trust and diagnostics for when they miss.
Where the Repeats Come From
CI is the best case. Prompts live in the repository and change only when someone edits one, so a run that doesn’t touch them should be almost all hits. The cache also makes the suite deterministic, which matters more than the money: a cached answer is the same answer on every run, so a flaky assertion about model output stops flaking. Add a strict mode where a miss fails the build instead of calling the provider (the way a locked VCR cassette behaves in Ruby HTTP tests) and CI no longer needs an API key at all. Recording once and replaying without servers is the same idea applied to an agent’s MCP traffic in CI.
Evals come next, with a catch. When you’re iterating on the grader, the rubric or the report, the model outputs don’t need regenerating, and caching turns a slow, expensive rerun into a quick one. When the eval measures sampling variance (five samples per prompt, say), the key has to include a sample index. Otherwise the cache collapses five samples into one and reports a suspiciously consistent model.
Agent retries are the case people underrate. A long run that dies at step 37 can restart and replay steps 1 to 36 from cache in seconds, as long as each step’s request comes out identical, which it does when the tools return the same results. The cache makes replay cheap. It doesn’t make it safe: if step 12 asked for send_invoice, the cached response asks again, and the agent will send it again unless its runtime keeps its own record of which tool calls already ran.
Then there are background jobs delivered twice by an at-least-once queue, and development loops, which repeat more than anyone admits.
Live chat sits at the other end. Past the first turn, the history makes each request unique; the first turn repeats in meaning (“what are your hours”, “when do you open”) but rarely in bytes. That’s where semantic caching tempts people. It embeds the request and serves a stored answer when a new one lands within a similarity threshold, which raises the hit rate and the odds of answering a different question. “Cancel my order” and “don’t cancel my order” share almost every token and can sit close together in embedding space. If you want a semantic tier anyway, make it opt-in per route, limit it to a short allowlist of FAQ-style routes with no account data, set the threshold high, and log every semantic hit next to the question it matched, for a person to read.
The Key Is the Product
An exact-match cache is a hash table, so the whole design lives in the key. Include every field that can change the output; leave out anything that can’t. A config for the proposed proxy might look like this:
listen: 127.0.0.1:8787
store: ./llm-cache.db # SQLite, one row per key
key:
fields: [model, system, messages, tools, tool_choice, temperature,
top_p, max_tokens, stop, response_format, seed]
namespace: header:X-Tenant # tenants never share entries
cache_when:
- header: "X-Cache: on" # explicit opt-in
- temperature: 0
never_cache:
tool_calls: [send_invoice, charge_card, delete_*]
provider_side_history: true
retention: 30d
encryption: per-tenant-key
Model ID means the exact snapshot. A floating alias that the provider repoints can’t be cached safely unless the config pins it, and each entry should record which model produced it so one command can drop everything from a retired snapshot. Requests that keep conversation state on the provider’s side, where the body carries a reference to an earlier response instead of the full history, can’t be keyed by their body at all and should pass straight through.
Temperature zero is the default trigger, and it’s worth being plain about what it buys. Temperature zero narrows the output without pinning it. Providers batch requests on shared hardware, and floating-point sums can come out differently depending on what else is in the batch. The serving stack also changes underneath you. A cache doesn’t need determinism, though. Any answer the model gave once for an exact input is an answer it could give again, so a cache freezes one valid sample. At higher temperatures that’s fine for tests and wrong for a feature that wants variety, which is why anything above zero needs the explicit opt-in.
The most common silent miss is a timestamp in the system prompt. “Current time: 2026-10-05T09:14:03Z” changes every second, so every request is new, and because it sits near the top of the prompt it breaks the provider’s prefix cache from that point on as well. Round it to the day or move it to the end. Better, make prompt assembly deterministic (same inputs in, same bytes out), which is a strong argument for treating context assembly as a database query instead of string concatenation. The proxy should help you find these misses: store a hash per field and per message with each entry, and an explain-miss command can find the nearest stored request and print the field that differed.
$ llmcache explain-miss req_7f3a
nearest 9c1e04 (stored 2026-10-04 22:10, 41 hits)
differs system prompt at offset 212
stored "Current time: 2026-10-04T22:10:07Z"
now "Current time: 2026-10-05T09:14:03Z"
Streaming needs care in both directions. Store only streams that finished cleanly, because a stream cut off by a dropped connection looks like a short answer and would be served as one until it expired. On a hit, replay in the caller’s shape (Anthropic’s named events from message_start to message_stop, or OpenAI’s chat completion chunks ending in data: [DONE]), split into token-sized deltas, since some client code breaks on one giant chunk. A small delay between chunks is realistic for UI tests; CI turns it off.
Tool calls are where caching can do damage. Replaying a response that asks for a read-only lookup is harmless. Replaying one that asks for a charge or a delete repeats a decision whose consequences already happened once, so those tools go on the never-cache list. MCP’s readOnlyHint and destructiveHint annotations can seed that list, but they’re hints from the server’s author and nothing more. Caching the tool call itself, as opposed to the model’s decision to make it, is a separate problem: the new MCP spec caches tool lists but not tool calls.
Privacy is the last part and the least optional. The cache is a database of every prompt and answer that passed through it, which makes it the most sensitive file in the build. Encrypt it at rest with a key per tenant, so deleting a tenant, or honoring a deletion request, means destroying a key. Keep a retention limit. Keep tenants apart even when their requests are byte-identical, because a shared cache leaks through timing alone: an instant answer tells one tenant that someone else already asked.
Version 0.1
The first version does exact-match caching for the OpenAI and Anthropic request shapes, streaming replay, storage in a single SQLite file or an embedded key-value store, and a savings report. The report is the feature people will screenshot, so it should be honest about where savings come from. The numbers below are invented; the shape is the point.
$ llmcache report --since 7d
route requests hits hit rate tokens avoided est. saved
ci/integration 8,412 7,988 95.0% 61.2M $214.20
evals/nightly 2,100 1,474 70.2% 18.9M $66.15
agent/replay 930 611 65.7% 4.4M $15.40
app/chat 14,233 38 0.3% 0.1M $0.35
prices from prices.yaml, edited 2026-09-30
Prices come from a config file because they change. The report prints to a terminal and exports CSV, and that’s the whole interface.
It refuses two things. Semantic caching stays off by default, behind the per-route opt-in described above. There’s no dashboard either, because the moment there is one, someone asks for a hosted version of the most sensitive file in the build. Picking which model answers is a different job, for a rules-based router.
Start with the test suite. Nothing repeats itself like CI.