Below you will find pages that utilize the taxonomy term “Cost Optimization”
Posts
LLM Response Caching Pays Off in CI, Evals and Agent Retries, Not in Chat
A repository has 300 integration tests that each send a prompt to a model. The suite runs on every push to every branch, then again in the merge queue. Most of those runs don’t touch a prompt (the diff was a stylesheet or a migration), so the model receives the same 300 requests it saw an hour earlier, writes roughly the same 300 answers, and the provider bills every one. Nobody chose that. It’s what happens when a test calls a live API.