Mapping an Unfamiliar API by Following the IDs Between Its Endpoints
You inherit an integration with a partner API that has 140 endpoints, documented in alphabetical order. GET /accounts sits next to GET /adjustments, and POST /refunds is a long scroll away from the GET /charges/{id} it depends on. Nobody wrote down that a refund hangs off a charge, which hangs off an order, which may or may not have a customer. You find out the slow way: call an endpoint, copy an ID out of the response, paste it into the next call, and keep a diagram in a notebook that’s out of date by Thursday.
The API already knows its own object graph. It just doesn’t say so. An id returned by one endpoint turns up as a path parameter, a query parameter or a nested field in another, so a tool that records every value it sees, and notices where those values reappear, can draw the entity map for you. The same pass can mark fields that behave like enums and endpoints that answer even though the docs never mention them. The output worth building is a readable map: a diagram plus sentences like “every order has exactly one customer” that a new engineer can read on day one. Plenty of tools touch this problem. None of them hands you the map.
The Closest Relatives Are Fuzzers
Microsoft Research’s RESTler, from 2019, is the nearest thing to following the IDs. It reads an OpenAPI spec, infers which request produces a value that another request consumes (create a post, take its id, use it in the GET and the DELETE), and builds request sequences out of those dependencies. It does this to find bugs, so it creates, mutates and deletes on purpose, and what it reports is failures. Schemathesis runs property-based tests from OpenAPI and GraphQL schemas and can chain calls through OpenAPI links, the part of a spec that says “the id in this response feeds that operation”. Links are optional and plenty of specs skip them, which tells you how much of the graph lives only in people’s heads. Both tools generate requests from a schema in order to break things. That’s a fine job (the same instinct drives generating bad arguments from an MCP tool schema), but an explorer needs the opposite temperament.
Two more projects came at the problem from the side. Akita Software inferred API specs by watching traffic, and Postman bought the company in 2023. Assetnote’s Kiterunner, released in 2021, finds routes for security testing by trying paths from wordlists built out of real API specs. Postman, Insomnia and Bruno, used as clients, are excellent at sending a request you already know how to write. Each of these produces something useful. The map as a finished product, the document you’d hand a new hire, is the gap.
How the Crawl Works
Seed the crawl from an OpenAPI file when one exists, or from a base URL and a token when it doesn’t. The spec gives you path templates to fill; without one, you start from whatever the root returns and from links found along the way. By default the crawler sends only methods RFC 9110 defines as safe: GET and HEAD for data, OPTIONS to read the Allow header.
Every response is stored, and every scalar in it goes into an index keyed by value, with the JSON path, the endpoint and a guessed type and format alongside. That index is the whole trick. When order.customer_id holds 8812 and the spec has a /customers/{id} template, the crawler proposes an edge and tests it with a confirmation probe: fetch /customers/8812 and see whether a customer comes back. One hit means little. Forty out of forty, drawn from different orders, with field names that line up, is an edge you can draw with a solid line. It’s the same inference problem as guessing the joins between files in a folder: find a column whose values sit inside another column’s values, then check that the overlap means something.
Cardinality and optionality fall out of the same samples. If 1,900 sampled orders carry 640 distinct customer IDs, customers have many orders. If a field appears in the schema and in 3% of records, it’s a different field from the one the docs describe, and the field completeness audit shows why the records that have it are rarely a random 3%. Enums show up as fields with few distinct values that look like identifiers (paid, on_hold) and, more telling, a count that stopped growing. Record when the last new value appeared. “Five values in 1,900 samples, none new after sample 240” is evidence; “status is an enum” is a guess with good posture.
Pagination detection decides how much you crawl. Look for next URLs in bodies, a Link header with rel="next", cursor and offset parameters, and stop after a few pages per collection, because the goal is a sample. Offset pages also shift when rows are inserted mid-crawl (one of several ways offset pagination misbehaves), so the same record can land in the sample twice and skew the counts. Dedupe by ID before computing anything.
Undocumented endpoints come from places the spec doesn’t control: hypermedia links in responses, next URLs that point somewhere unexpected, Location headers, an Allow header that lists a DELETE nobody documented, and string literals in the JavaScript bundle of the vendor’s own web app, which often calls endpoints the public docs leave out. Anything seen in the wild and missing from the spec gets its own section in the output. A run of the proposed crawler against a partner sandbox might print this:
$ idcrawl --spec partner-openapi.yaml --base https://sandbox.partner.example --bearer-env PARTNER_TOKEN
61 of 64 documented GET operations crawled, 2,317 requests, 0 writes, 3 returned 403
erDiagram
CUSTOMER ||--o{ ORDER : customer_id
ORDER ||--|{ LINE_ITEM : order_id
LINE_ITEM }o--|| PRODUCT : product_id
ORDER ||--o{ CHARGE : order_id
CHARGE ||--o{ REFUND : charge_id
Every order has exactly one customer (order.customer_id, 40 of 40 probes confirmed).
A customer has 1 to 212 orders in the sample, median 3.
3% of sampled orders have no shipping_address.
order.status behaves like an enum: 5 values in 1,900 samples, none new after sample 240.
invoice.account_ref looks like a customer ID, but product IDs resolve there too: unconfirmed.
GET /orders/{id}/audit-log is not in the spec. Found in app.js, answers 200.
The Mermaid block pastes straight into a README, and a JSON graph file written next to it gives other tools the same edges with their evidence attached.
Where It Goes Wrong
Safety comes first, and part of it is legal. Point the crawler only at APIs you own or have written permission to explore, and prefer a sandbox; most serious APIs offer one, and exploring a sandbox costs nothing when the crawler misbehaves. GET isn’t safe everywhere, either. RFC 9110 calls it safe on the server’s behalf, and servers break that promise all the time: a GET that marks notifications as read, an export endpoint that kicks off an expensive job, a /logout that kills the token the crawl depends on, a “preview” link that sends the email it previews. Ship a denylist of suspicious path words and an opt-in list for anything else that might mutate. Treat a 429 with Retry-After as an instruction, and give the whole crawl a request budget that works as a hard stop.
Sequential integer IDs are the next trap. With UUIDs, a value that appears in two places is almost certainly the same object, and prefixed IDs in Stripe’s style (cus_, ch_) carry their type with them. The integer 42 is an order, a customer, a product and a warehouse at once. A probe for /customers/42 succeeds whether or not the 42 you took from an invoice meant a customer, so a 200 proves nothing on its own. The fix is a negative control: also probe the template with IDs taken from an unrelated collection. If product IDs resolve under /customers/{id} just as often, the endpoint accepts any small integer and the match is noise. Supporting evidence helps too, such as the returned customer listing the order you started from. Report the evidence with every edge, and draw unconfirmed ones dotted.
Sampling bias is quieter. An enum looks closed after 200 samples and isn’t; the sixth status value shows up at month end, or only for accounts on a plan your test account isn’t on. Sandbox data is worse, since a seeding script made it and it has none of the half-migrated records that make production interesting. The map has to say what it saw and how much of it, on every line.
Scopes hide whole regions of the graph. A read-only key gets a 403 on refunds, and unless the map marks that as a hole, it shows orders that never have refunds. Run the crawl with two tokens of different privilege and diff the maps; you get a rough picture of the permission model for free. Some collections never end, too. A calendar endpoint will hand you a next link for every future day, so cap pages per collection and stop when a cursor repeats.
A First Version, Then a Watch Mode
Version 0.1 should be small enough to finish. It’s seeded from OpenAPI, sends GET, HEAD and OPTIONS only, supports one auth method (a bearer token or an API key header), and writes Mermaid plus the JSON graph. It refuses writes outright, with no flag to turn them on yet. It refuses GraphQL, which is a separate problem: introspection, where it’s enabled, hands you the schema, and the relationships live in the type system. It also refuses OAuth flows, mTLS and signed requests such as AWS SigV4; each is a week of work that teaches you nothing about mapping.
The paid version is the same crawl on a schedule. Run it nightly against the APIs you depend on, diff the inferred map, and alert on changes: a new field, a vanished field, a new enum value, an endpoint that started returning 404. If you own the API, inferring the contract from your own traffic gets you the same signal passively, with far more samples. For a partner’s API you can’t instrument, active probing is the only way to see a change before your integration does. Teams that have been broken by a silent new enum value remember the cost, and they’re the buyers.
The graph is already in the responses.