The Best Small Infrastructure Tools Reduce Data Near the Source Instead of Storing More
A gigabyte of disk costs almost nothing. The same gigabyte sent to a log vendor that prices by ingestion costs real money, and pasted into a model’s context window it costs money and answer quality at the same time. Storage is cheap per byte. Everything that touches the bytes afterward is not: ingestion fees, query time, bandwidth on a thin link, a context window that holds only so much, and the attention of whoever has to read the result. Observability is the oldest place this shows up, and agents have just made it louder.
That’s why the small infrastructure tools I’d build first are reducers, not databases. Two brainstorm lists of tool ideas fed this batch of posts. When I collapsed them, the ideas that kept their shape were the ones that take a firehose near where it starts and keep the information instead of the bytes. A lightweight database is a crowded idea, because storing things is the part of the problem everyone already knows how to sell. Reduction near the source is less crowded, and I think it’s the better bet.
Every reducer has to answer one question before it ships: what must survive? State that and you have a contract. This answer stays exact, these events stay whole, this error bound holds, and the rest may go. If you can’t state it, you’ve built a lossy compressor with no contract, and the first time it throws away the thing someone needed, nobody can say whether that was a bug. A contract fits in a few lines. This is a sketch for a log pipeline (illustrative; no such format exists):
survives:
exact: [error_count, request_count]
whole: ["level >= error", "template never seen before"]
bounded: { p99_latency_ms: "1% relative error", unique_users: "1% standard error" }
drops: [debug lines, repeated templates, unchanged readings]
raw_kept: 48h
report: counts_and_names # every output says what was cut
None of this is new, which is a good sign. RRDtool, from 1999, keeps metrics in round-robin archives with consolidation functions (average, min, max, last), so the file never grows and old data gets coarser on a fixed schedule. Process historians use deadband filtering and swinging-door compression, popularized by OSIsoft PI, to store a reading only when it leaves a tolerance around what the last stored ones predicted. Both are explicit about what survives. On the commercial side, Cribl built a large company on routing and reducing telemetry before it reaches expensive backends. Reduction was a business before anyone put a model in the loop.
Seven Reducers, One Question
Start with the newest case. An agent calls a search tool, gets back a hundred hits with dozens of fields each, and needed three. What survives is the answer-bearing fields, the counts and a way back to the rest. A deterministic reducer between server and model applies per-tool rules and stashes whatever it cuts behind an expand call, so a wrong guess about what mattered costs one more call.
Ordinary API clients have the same problem, and the fix is older than MCP. Google’s APIs take a fields parameter, JSON:API has sparse fieldsets, OData has $select, and GraphQL is built on field selection. What’s missing is the same thing for APIs that offer none of those, which a proxy can supply by keeping the few fields the client reads.
Logs are mostly a handful of templates repeated with different numbers in the blanks. Learn the templates (Drain is the usual algorithm) and a reducer can ship counts per template and new templates whole, instead of every line. Vendors sell versions of this. Grafana Cloud’s Adaptive Logs finds low-value patterns so they can be dropped or sampled, Edge Delta processes telemetry at the edge, and Datadog’s Logging without Limits separates ingestion from indexing. Reducing logs locally before shipping them is the same move for a team that doesn’t want a vendor in the middle, and the Logs part of Precomputing does a version of it: it learns templates with Drain, keeps raw lines on the machine for 48 hours and sends reduced answers upstream.
Old metrics don’t need the resolution new ones do. Thanos downsamples to 5-minute and 1-hour resolution on the same logic. A telemetry database that forgets on purpose keeps rollups and anomalies instead of raw rows, and a rollup table is the cheapest reducer there is: the cookieless analytics idea in Nine Apps Worth Coding keeps SQLite rollups instead of raw hits.
Some questions can be answered from a summary with a stated error. Redis counts distinct values in a HyperLogLog of about 12 KB per key, with a standard error of about 0.81%, however many events went in. An embedded sketch database is that idea with error bars on every answer, so the reduction shows in the output. Here the contract is a number: right to within this much.
An agent’s accumulated memory and history grow without limit while its window doesn’t. What survives is whatever the next step needs, ranked by usefulness under a token budget. Context window management as a database frames it as asking for 12,000 tokens and getting the most useful ones.
Last comes the cost of all the others. Reduction makes systems cheaper and harder to inspect, so something has to keep what remains legible. A flight recorder in one portable file, the way HAR did it for HTTP, records what the agent did and what each reducer dropped, so a bad answer can be traced back to the cut that caused it.
The Edge Makes the Case Plainly
A device on a thin or intermittent link can’t ship everything, so it has to decide what’s worth sending. AltSql (a project from this site’s publisher) is built on that. Devices keep records in their own flash behind a key-value interface and go on making decisions offline, the gateway database stores the same records byte for byte and answers SQL over them, and the two sides reconcile when the link returns. The device-side core compiles to about 15 KB of code for a microcontroller. AltSql Ask puts one SQL question to a whole fleet and brings back only the answers, so the question travels to the data instead of the data traveling to the question. It’s alpha, with the hardware simulated so far.
Reduction Deletes Things, So Plan for It
The objection is correct: reduction deletes. Somebody will want the line you dropped, the row you rolled up or the field you projected away, usually in the middle of an incident. Four habits make that survivable.
Keep raw data locally for a short window, long enough for the question to arrive. Keep anomalies whole, so the reducer compresses the ordinary and lets the unusual through untouched. Make every reducer report what it dropped, counted and named, in the place the consumer reads. And measure the misses as well as the savings, such as how often someone went back for the raw data or how often a model called an expand tool after a reduction.
Savings alone prove nothing, since a reducer that drops everything saves 100%. The measure that matters is whether answers from the reduced data match answers from the raw data, on a sample you keep around to check. A reducer that can’t pass that check is only deleting.
Reducers also fit the shape that lets small tools survive. They sit in a pipe, read one config and need no cluster. One binary, one file, no daemon is how SQLite and nginx earned their place, and a reducer suits that shape better than yet another database does, because it replaces a bill instead of adding a system.
Say what survives. Drop the rest.