Finding API Endpoints, Tables and Flags Nobody Uses Takes Code and Traffic Together
A team wants to delete GET /v1/invoices/{id}/legacy-pdf. A search across the main repo finds no caller, and the gateway logs show zero hits in the last 30 days. The route goes out in a cleanup PR. On the first business day of the next quarter, a partner’s reconciliation job starts failing: it calls that endpoint four times a year, and nobody on the team knew the partner was still there.
Both checks were honest. Neither was enough.
That’s the trap in hunting for the dead parts of a system. Static analysis can’t prove something is dead, because it can’t see names assembled from strings at runtime or callers that live outside the code it was given. Runtime evidence can’t prove it either; a quiet month says nothing about a job that runs in January, and a mobile app released in 2024 that can’t be force-updated still calls whatever it called on launch day. Used alone, each kind of evidence produces a list that looks official and is wrong in exactly the cases that hurt. The tool worth building joins them. It keeps an inventory of everything that might be dead (routes, tables, columns, config keys, feature flags, environment variables), gathers static references from every repo plus runtime hits over a window you declare per item type, and scores each item. Then it walks you through removal in steps that fail loudly and can be undone.
What Already Exists
Each item type has decent tooling of its own. Knip finds unused files, exports and dependencies in JavaScript and TypeScript projects. Vulture does dead code for Python and is upfront about its limits: every finding carries a confidence percentage, because Python is too dynamic for certainty. Go’s deadcode tool, announced late 2023, reports functions that can’t be reached from main. Uber’s Piranha goes past detection and rewrites code to delete stale feature flags along with the branch that no longer runs. Flag platforms such as LaunchDarkly and Unleash can mark flags as stale or show where in the code a flag is referenced.
On the data side, Postgres keeps per-table counters in pg_stat_user_tables (seq_scan and idx_scan for reads; n_tup_ins, n_tup_upd and n_tup_del for writes), while pg_stat_user_indexes shows idx_scan = 0 for indexes nothing uses. That last query is the classic way to find indexes worth dropping. And every API gateway or load balancer writes access logs that say who called what.
None of these talk to each other, and none of them reason about time. Knip has never seen your traffic. A gateway log can’t know about the cron job in the infrastructure repo that calls an endpoint on December 31st. Postgres counters know a table was read but not by whom, and they go back to zero when someone resets statistics or the server crashes, so one reading of idx_scan = 0 could mean “unused for a year” or “unused since Tuesday”. The traffic you’d capture to infer an API’s real contract is the same traffic this tool needs, read with a different question.
One Report, Ranked by Confidence
Start with the inventory, because you can’t score what you haven’t listed. Routes come from framework introspection (rails routes, or the OpenAPI document FastAPI generates) or from a spec you trust. Tables and columns come from information_schema. Config keys come from the config files, flags from the flag service’s API, and environment variables from deployment manifests plus the code that reads them.
Static evidence means references across every repo the organization owns: services, infrastructure, data pipelines, the BI project, both mobile apps. Searching one repo is how the partner story above happens. Runtime evidence means gateway logs grouped by route template (/v1/invoices/{id}/legacy-pdf, never the raw path with an ID in it), daily snapshots of the Postgres counters written into the tool’s own table, and evaluation counts from the flag service. Config keys and environment variables usually have no runtime evidence at all, since almost nothing logs a config read.
The time window has to be declared per item type, and it has to be long. Ninety days is fair for endpoints behind an app. Anything that might run yearly (tax exports, annual statements, certificate renewals, the year-end board report) needs about 400 days, a full year plus margin. An item scores high only when it has no static references and no runtime hits across an unbroken window. Gaps in the evidence, like a week of missing logs or a stats reset, cap the score instead of counting as silence.
A run of the proposed tool might print something like this:
$ unused report --since 400d
ITEM LAST SEEN STATIC RUNTIME (WINDOW) CONF NEXT STEP
GET /v1/invoices/{id}/legacy-pdf 2026-03-14 0 0 hits (90d) 0.86 brownout, 410 for 1h
flag checkout_v2_rollout 2026-10-04 3 always true (30d) 0.90 delete flag, keep on-path
table audit_events_v1 never read 0 writes only (400d) 0.71 find the writer
env LEGACY_SMTP_HOST unknown 0 no runtime signal 0.40 assign an owner
column users.fax_number unknown 2 no column stats 0.30 check BI queries
Two rows deserve a closer look. The flag was evaluated yesterday, so it counts as used, yet it has returned true for every user for a month: dead weight with traffic. The table is the opposite case. Something still writes to it, so it looks busy on every dashboard, while nothing has read it in over a year. A simple “zero hits” query misses both.
Where It Gets Hard
Dynamic references come first, because they make static evidence wrong in both directions. Python’s getattr(handlers, name), Node’s process.env[key], a table name built as f"events_{year}", an ORM that maps the Person model to a people table so the string “people” never appears in your code. A scanner that doesn’t know the framework’s naming rules will call half the schema unreferenced. One that finds a dynamic access should lower its confidence for every name the expression could produce; skipping the line is how false positives get in.
Columns are worse. Postgres keeps no per-column counters, so column usage has to come from parsing queries, either from logs or from the normalized statements in pg_stat_statements. An ORM that selects every mapped column marks fax_number as read on every user lookup, whether or not any code touches the attribute. At the database layer, a column the ORM still maps looks exactly as alive as the primary key.
Then come the consumers you can’t see. Hyrum’s Law says that with enough users, every observable behavior will be depended on by somebody, which is how every accident in your API becomes a contract. Old mobile app versions keep calling what they called on release day. Partners call from systems you’ll never scan. BI tools often query tables directly, sometimes through a read replica, and a replica keeps its own statistics, so the primary’s counters never see those reads. Snapshot every instance. Treat “no evidence” for a table analysts can reach as weaker than “no evidence” for one only your service can touch, and assume there’s a cron job on somebody’s laptop.
Seasonal callers need their own handling. Three hits spaced ninety days apart is a schedule, and a report that only prints “last seen 205 days ago” hides it. Show the hit history for low-traffic items, and flag a regular interval as a likely periodic caller with a predicted next date.
The counters need care too. Treat the Postgres stats the way Prometheus treats counters: store snapshots, read any decrease as a reset, and sum the deltas instead of trusting the latest value. A nightly ETL job that copies every table into a warehouse makes every table look read, so the tool has to exclude known bulk readers. That’s easy when the job has its own database role and miserable when it shares the application’s.
Finally, visibility. The search has to cover repos the person running it may not be allowed to read, which turns a CLI into an org-wide crawler holding credentials. That’s an access decision as much as a technical one, and it’s where a hosted version would earn money: platform teams at companies old enough to run three generations of APIs at once.
Removal in Reversible Steps
A high score is a reason to start a process. For an endpoint, the standard retirement sequence for integrators applies even when you believe nobody is left:
- Mark it deprecated and send a
Sunsetheader with the planned date. - Log every call with the caller’s identity, and alert the owner on each one.
- Brown it out: return 410 Gone for an hour, at a time you pick, and see who shouts. It’s a scream test run while people are awake.
- Remove the handler, but keep the route answering 410 for a while.
Tables get the database equivalent. Revoke grants from application roles, or rename the table, then wait before dropping it; both undo in one statement, while a DROP TABLE undoes only from a backup. Flags get Piranha’s treatment, with the check and the dead branch deleted together. Removing an endpoint is the most obvious breaking change there is, and a confidence score doesn’t change that.
Version 0.1 should be small: one framework’s routes plus its access logs, daily Postgres stats snapshots, a scan for environment variables and config keys across a list of repos, and the ranked report. Config is the weakest evidence, so it helps that a linter reading env files, YAML, TOML and CI config together finds keys defined in one place and read nowhere, which is this problem seen from the files. The inverse question, which features break if a dependency disappears, runs on the same inventory.
What v0.1 refuses to do is delete anything. Its output is a probability, and a deletion needs an owner whose name is on the pager.
A quiet log is a hint. Treat it like one.