Recent Posts
A Tiny ETL Binary Competes With curl, jq and SQLite in a Cron Job, So Build It That Small
The real competitor is a shell script in a crontab, and it usually looks like this:
curl -s "https://api.example.com/v1/launches?limit=100" \
| jq -r '.data[] | [.id, .name, .net] | @csv' \
| sqlite3 -csv launches.db ".import /dev/stdin launches"
It works on the day you write it. Then the API answers 429 and curl pipes an error page into jq. Or the API has a second page, and the script never asks for it. When the job dies halfway, the rerun inserts the same rows again or trips over the primary key, depending on how the table was made. A field that starts arriving as a string goes unnoticed until a chart looks wrong. Retries, backoff, pagination, incremental state, idempotent writes and schema drift: that’s the list, and shell scripts get every item on it wrong in predictable ways. (To be fair, curl --retry covers the first two.)
An AI Agent Flight Recorder Belongs in One Portable File, the Way HAR Did It for HTTP
An agent edits the wrong config file, then spends forty minutes trying to repair its own repair. The user files a bug with a screenshot of the last message. The maintainer asks for logs, and the logs live in four places: the provider’s usage page, the tool server’s stdout, the framework’s debug output, and a terminal scrollback that closed with the window. Nobody can say what the agent saw on turn six.
Benchmarking MCP Servers and Stateful API Workflows: Latency, Throughput and Tokens per Call
Point wrk at an MCP server and you get a clean report: every response a 200. Some of those 200s carry isError: true results, and nothing in the output says what the tool definitions cost the model on every turn. The tool is measuring requests per second. An agent runs a workflow, and what the server costs it is latency per call times calls per task, plus the tokens each response and each definition puts into its context.
Diagnose a Slow API From One Request: DNS, Connect, TLS, Server Wait and Transfer
“The API is slow” starts a hunt. Somebody opens the APM dashboard, somebody else greps the load balancer logs, a third person checks whether the database is on fire. An hour later the team has three theories and no measurement. One request, timed phase by phase, answers the first question in under a second: whether the time went to name lookup, the TCP connect, the TLS handshake, the server or the transfer of the body.
Finding API Endpoints, Tables and Flags Nobody Uses Takes Code and Traffic Together
A team wants to delete GET /v1/invoices/{id}/legacy-pdf. A search across the main repo finds no caller, and the gateway logs show zero hits in the last 30 days. The route goes out in a cleanup PR. On the first business day of the next quarter, a partner’s reconciliation job starts failing: it calls that endpoint four times a year, and nobody on the team knew the partner was still there.
Fuzzing MCP Servers: Generate Bad Arguments From the Tool Schema and Watch What Breaks
You test an MCP server by chatting with it. Ask for the open bugs, get the open bugs, ship it. The model never sends limit: "ten", an empty path or a 200 KB query while you’re watching, so those paths stay dark until a real session hits one. Say it’s the limit. The handler throws, the framework wraps the exception in a 40 KB stack trace, and the model reads all of it, adjusts, and retries with the same bug in a new shape.
Give Any API a History: Poll It, Hash It and Query Old Versions With SQL
Ask a launch schedule API when a rocket flies and you get one date. Ask what the date was last Tuesday, or how many times it has moved, and there’s no endpoint for that. Most APIs describe the present. Prices, timetables, government datasets and status pages all change in place, and the old value is gone the moment the new one is written. That’s a pity, because launch dates slip so often that the list of changes in a launch schedule is more interesting than the current date.
Infer Your API's Real Contract From Traffic, Then Diff It Against the Docs
Your OpenAPI file marks shipping_address as required on GET /orders/{id}. After Tuesday’s deploy, about one response in 300 leaves it out, because a new code path for orders created by the import job skips the field. The docs still say required. The tests still pass, since nobody wrote one for import-job orders. A client with a strict deserializer starts throwing on 0.3% of order pages, and the first report you get is “sometimes the order screen is blank”.
Keeping Third-Party API Responses in SQLite Gets You a Cache, an Offline Mode and a History
The first version of an API cache is a dictionary with a timeout. The second is Redis holding a JSON string under a key built from the URL. The third gets written after an incident, when somebody needs to know what the weather provider returned on Tuesday and the cache has already replaced it with Wednesday’s answer. Every app that depends on an outside API walks the same path: a cache, then a retry layer, then a debugging log, then a wish that it had kept the old responses.
LLM Response Caching Pays Off in CI, Evals and Agent Retries, Not in Chat
A repository has 300 integration tests that each send a prompt to a model. The suite runs on every push to every branch, then again in the merge queue. Most of those runs don’t touch a prompt (the diff was a stylesheet or a migration), so the model receives the same 300 requests it saw an hour earlier, writes roughly the same 300 answers, and the provider bills every one. Nobody chose that. It’s what happens when a test calls a live API.