Routing LLM Requests by Token Count, Privacy Tag and Cost Ceiling From One Config File
LLM routing usually starts as if statements. One service picks the cheap model for ticket summaries. Another hardcodes the strong model because it needs tool calling. A third has a comment reading “never send this to the cloud” directly above a fallback that sends it to the cloud when the local server times out. Changing a model, a price or a provider means editing all three.
The router worth building is a small one: ordered rules over properties you can compute before the request leaves, written like an nginx config in a single file, with no database and no UI. Input size, whether tools or images are present, a data classification tag, a cost ceiling for the single request, which upstreams are healthy. One binary on a cheap VPS reads the file, accepts the OpenAI request shape and picks where each request goes. The rule that justifies the exercise is that privacy routing fails closed.
This is a crowded category, and a newcomer should say so first. LiteLLM’s proxy, OpenRouter, Portkey, Kong’s and Envoy’s AI gateways and Cloudflare AI Gateway all sit between your code and the model providers, and between them they cover keys, budgets, fallbacks, caching and logs (the gateway section of this review goes through several). RouteLLM, from LMSYS in 2024, goes the other way and learns when a cheaper model is good enough for a given query. Any claim about what these products lack will age fast, so check each one’s policy hooks before building, and if a gateway you already run can express your rules, build versus buy answers itself. The gap worth betting on is size and legibility: a config one person can read on a single screen, and a command that says why a request went where it did.
Rules on Things You Can Count Before Sending
Two of the obvious predicates are harder than they look. Start with token count. Providers tokenize differently, and while some publish a tokenizer library or a counting endpoint, others give you neither. A counting endpoint is also a call to a cloud provider, which rules it out for exactly the requests that mustn’t reach one. The workable default is local: count bytes, divide by a conservative ratio from config, add a margin, and widen the margin for code, JSON and non-Latin text. When the estimate lands near a threshold, round toward the larger context window. Misrouting up costs a few cents; misrouting down fails the request.
Expected output can only be bounded, and max_tokens is the bound. The cost check uses it: estimated input cost plus max_tokens times the output price. A client that sets max_tokens to the maximum on every call will look expensive, and the dry run will show why. One of the two big API shapes requires that field and the other doesn’t, so the router needs a default for callers who omit it.
“If it’s coding” needs a classifier, and a classifier is either a model call (extra latency, extra cost, and a privacy problem when the thing being classified is confidential) or a pile of heuristics that fails on the first odd input. Let the caller say. A header like X-Task: coding costs nothing, and the caller knows. The router acts on tags; it doesn’t infer them. Here’s a sketch of the file, with invented syntax.
upstream small { api openai; model small-model; price 0.25 1.00; } # per million tokens, example values
upstream big { api anthropic; model big-model; price 4.00 20.00; }
upstream local { api openai; url http://127.0.0.1:11434/v1;
model local-model; price 0 0; allow confidential; }
class { # the key sets the class; a header can raise it, never lower it
key support-bot = confidential;
header X-Data-Class;
}
route {
ceiling 0.05; # worst-case dollars per request, a veto like class
if class = confidential { use local; }
if tools or images { use big, small; }
if task = coding { use big; }
if tokens < 4000 { use small, big; }
default { use big, small; }
}
Privacy Is a Filter, Not a Rule
Rules pick an upstream and the first match wins, so they shouldn’t be the only thing between a confidential request and a cloud API. Rules get reordered. Somebody adds a rule above the confidential one at five on a Friday. So split the work: rules produce an ordered candidate list, then filters veto every candidate that isn’t allowed to see the request. If nothing survives, the router sends nothing anywhere and returns an error naming the class and the rule. A confidential request that matches the tools rule by mistake fails loudly instead of leaking quietly. The cost ceiling is the second veto, applied the same way.
The class has to come from somewhere harder to forget than a header. Attach it to keys: the support bot’s key is always confidential, whatever it sends. A header can raise the class for one request, never lower it. Fail-closed has corollaries. A local server that’s down isn’t an outage to route around, so no cloud fallback ever applies to that class. The cost log for confidential traffic holds counts and prices, no content. And the dry run takes metadata only, so you can debug a routing decision without pasting a prompt into a terminal.
Translation Is Where the Bugs Live
Routing across providers means translating between request shapes, and that’s most of the code. A system prompt is a message with a role in one shape and a top-level field in the other. Tool definitions use different keys. A tool call’s arguments arrive as a JSON string in one and as an object in the other. Tool results are their own messages here and content blocks inside a user message there. Finish reasons have different names (tool_calls against tool_use, length against max_tokens). Streaming is worse: one family ends with a [DONE] sentinel, the other sends named events, and both stream tool arguments as fragments of partial JSON that must be reassembled in order. A translator that gets one of these wrong passes every simple test, then breaks on the first agent that makes three tool calls at once.
Fallback has its own rules. You can fall back on a refused connection, a 429, or a 5xx that arrives before the first byte. Once the first streamed token reaches the client you can’t switch providers; the router can only end the stream with an error and let the client retry. A read timeout is unsafe to retry elsewhere too, because the first provider may still be generating, and billing. Log both attempts. Give each upstream a circuit breaker: after a run of failures stop sending, wait, probe with one request.
Prices change, and so do model snapshots, so the price table is config with an effective date, and every log line records which table it used. The log is append-only JSON lines: time, request ID, rule, upstream, estimated input tokens, the count the provider reported, cost. Estimate and actual side by side let a report command say how far off the estimator runs and what margin it needs.
Then comes the feature that earns trust. llmroute explain sample.json prints each rule in order, which one matched and why, the worst-case cost, and the candidates left after the filters. That property is what BareProxy is built around: its explain names the rule a URL matches, and its plan lists which requests would change hands under a new config. BareProxy is a web server and reverse proxy that doesn’t route LLM traffic, so the code doesn’t transfer, only the idea. A model router needs both commands, because a price change is exactly the edit you want previewed against last week’s traffic before it ships.
What v0.1 Does and Refuses
Version 0.1 accepts OpenAI-shaped requests and talks to OpenAI-shaped and Anthropic-shaped upstreams. It does ordered rules on the token estimate, tools or images, and task and class tags, with the two filters, fallback chains with breakers, the cost log and explain. It refuses a dashboard, a database, learned or semantic routing, and response caching. The first feature request will be a dashboard, and the answer is the log file: JSON lines, and jq makes a decent dashboard. Caching has its own keys and its own privacy risks, covered in the response cache post. A small MCP proxy is the sibling for agent traffic, and a ceiling per request is only half the answer for unattended agents; cron for agents puts a ceiling on the whole run.
Write the privacy rule first. Everything else is tuning.