Make for APIs: Rerun Only the Steps Downstream of an Endpoint That Changed
A nightly job pulls a users endpoint and an orders endpoint, normalizes both, joins them on customer ID and renders a report. On most nights the upstream data hasn’t changed since the last run. The job doesn’t know that, so it downloads and parses everything, reruns every transform and writes the same report again. If the API bills per call, or one step takes twenty minutes, you pay for the same answer every night.
make solved this for files in the 1970s: a target gets rebuilt only when a prerequisite is newer than it is. HTTP has a built-in equivalent of the timestamp. A server that sends an ETag lets you ask “has this changed since the version I have?” with If-None-Match, and when nothing has, it answers 304 Not Modified with no body. Last-Modified and If-Modified-Since do the same job with dates. The tool worth building is make with URLs as prerequisites. It treats a validator (or a hash of the body, when the API sends none) the way make treats a modification time, and it treats each transform as a pure step whose output is cached by a hash of its inputs. Checking is nearly free. Rerunning is the expensive part, and it happens only downstream of something that moved.
If the provider offers webhooks, use them as the trigger (the post on webhooks versus polling covers the trade). This tool is for the many APIs that don’t, where you poll anyway.
Prior Art, Piece by Piece
A Makefile plus curl gets you the fetch half today. curl --etag-save and --etag-compare keep an ETag in a file, and -z sends a conditional request based on a file’s date. For two endpoints and one transform, that’s the right answer. Bazel and Nix cover the other half: they cache build actions by a hash of their inputs, so a step with identical inputs never runs twice. dbt gives you a DAG of SQL models with incremental models, but its inputs are tables in a warehouse, not responses from somebody else’s server. Dagster is asset-based, with freshness rules, but it’s a platform with a scheduler and a UI. Snakemake brings make-style rules to scientific workflows, and its inputs are files. Hugo caches remote data at build time through resources.GetRemote (see fetching a remote JSON API at build time), so a rebuild doesn’t hit the API again.
What’s missing from that list is a tool whose native input is an HTTP response with a validator, a fallback hash and a reason it can print. That’s a narrow gap.
A Config and a Plan
The config holds two kinds of target: fetch targets (URL, auth reference, change detection) and steps (declared inputs, a command, an output path). Everything else is derived. The command name below is a placeholder for a tool that doesn’t exist yet.
[fetch.users]
url = "https://api.example.com/v2/users"
auth = "env:EXAMPLE_TOKEN"
validate = "etag" # etag | last-modified | hash
[fetch.orders]
url = "https://api.example.com/v2/orders?status=open"
auth = "env:EXAMPLE_TOKEN"
validate = "hash" # this API sends no validators
ignore = ["/meta/generated_at"] # volatile fields stay out of the hash
[step.normalize_users]
needs = ["users"]
run = "jq -f normalize.jq {users} > {out}"
out = "build/users.json"
[step.join_orders]
needs = ["normalize_users", "orders"]
run = "python join.py {normalize_users} {orders} > {out}"
out = "build/joined.json"
[step.report]
needs = ["join_orders"]
run = "python report.py {join_orders} > {out}"
out = "build/report.html"
The command that earns its keep is plan, the equivalent of make -n. It makes the conditional requests (they’re cheap), runs no transforms, and prints what would rerun and why:
$ apimake plan
users 200, ETag W/"a1" -> W/"b7" changed
orders 304 Not Modified unchanged
normalize_users input users changed RERUN
join_orders input normalize_users changed RERUN
report input join_orders changed RERUN
3 of 5 targets would run. Nothing was executed.
Tomorrow users might come back 304 while orders changes, and the plan will show normalize_users skipped. If normalize_users does rerun but writes byte-identical output, join_orders is skipped too, because its input digest hasn’t moved. Build-system people call that early cutoff, and it’s what makes a changed ETag over unchanged content cheap. Inside a database the same idea goes by incremental view maintenance: update a result from the change instead of recomputing it.
A plan also makes the tool safe to edit. Change a transform, run plan, and see which outputs you just invalidated before cron finds out at 3 a.m. BareProxy’s plan command does this for a web server: it lists which requests a config change would move before the change goes live. A pipeline needs the same property.
Independent targets run in parallel. The fetch side (retries, pagination, writing rows somewhere) is what a tiny ETL binary would handle. This tool sits one layer up and decides whether any of that needs to run.
The State File Is Where It Goes Wrong
Where the state lives is the first design decision, and the obvious answer is wrong. If the state file stores the last ETag seen per URL and updates it on every fetch, a crash inside join_orders leaves the new ETag saved and the join undone. Tomorrow’s request gets a 304, everything looks current, and the report stays stale. make never has this bug, because it works out staleness from what’s on disk instead of remembering what happened last time. Copy that. Record, per step, a digest of the inputs it last built from (content hashes of its inputs plus its command line), call a step stale when the current digest differs, and write the record only after the output has been renamed into place. Fetch targets keep the body they received, because a 304 means “use your copy” and the copy has to exist. If it’s gone, the tool drops its validators and fetches unconditionally. Keep the older bodies and the state directory becomes a history of the API almost for free.
Many APIs send no validators, and some send ETags that change when the body didn’t (a different ETag per server behind a load balancer is a classic). For those you have to fetch to find out, so the savings are downstream only. Hash the canonical body, meaning parsed, volatile paths dropped, keys sorted. The ignore list exists because an API that stamps generated_at or a request ID into every response defeats any hash. A 200 whose canonical body hashes the same as last time counts as unchanged.
Collections are worse. The version of a collection is the hash of all its pages, because page one’s validator only describes page one. With offset pagination, rows inserted mid-crawl shift the later pages, so a row appears twice or never. Hash the assembled records, deduplicated and sorted by key, instead of the raw pages, or a shifted boundary looks like a change. Unless the API has an updated_since filter or a change feed, there’s no way around fetching every page.
Transforms have to be pure for any of this to work, and scripts often aren’t. A step that writes the current time into its output changes its own hash on every run, so everything below it reruns every night and the tool saves nothing. The cheap defense: the first time a step builds, run it twice, compare the output hashes and complain if they differ. A step that needs today’s date declares a daily clock input, and the plan prints “clock: new day” as the reason it reran. Secrets stay in environment variables referenced from the config, never in the config or the state file. When two tokens see different data, the cache key needs a label for the identity, never the token itself.
A failed fetch shouldn’t sink the run. Like make -k, the tool keeps building targets that don’t depend on the failure, retries 429 and 5xx responses with backoff, and exits non-zero at the end. Outputs are written to a temp file and renamed, so an interrupted step leaves yesterday’s output in place.
Scope for Version 0.1
Version 0.1 fetches HTTP targets and detects change by ETag, Last-Modified or content hash. It runs shell commands as steps with declared inputs. It has plan and run, with independent targets in parallel, and it keeps one local state directory. It doesn’t schedule (cron exists), it has no UI, it won’t distribute work across machines, and it leaves MCP for later.
MCP calls have the same shape, so the idea carries over. In the 2026-07-28 MCP revision, resources/read results can carry cache hints (ttlMs, cacheScope), which gives a resource a freshness signal. tools/call results carry none, so a tool call is the no-validator case: call it, hash the answer. Tool calls can have side effects, so a pipeline tool should refuse to rerun them casually.
Cron says when to look. The graph says whether anything needs doing.