Below you will find pages that utilize the taxonomy term “Privacy”
Routing LLM Requests by Token Count, Privacy Tag and Cost Ceiling From One Config File
LLM routing usually starts as if statements. One service picks the cheap model for ticket summaries. Another hardcodes the strong model because it needs tool calling. A third has a comment reading “never send this to the cloud” directly above a fallback that sends it to the cloud when the local server times out. Changing a model, a price or a provider means editing all three.
The router worth building is a small one: ordered rules over properties you can compute before the request leaves, written like an nginx config in a single file, with no database and no UI. Input size, whether tools or images are present, a data classification tag, a cost ceiling for the single request, which upstreams are healthy. One binary on a cheap VPS reads the file, accepts the OpenAI request shape and picks where each request goes. The rule that justifies the exercise is that privacy routing fails closed.