Diagnose a Slow API From One Request: DNS, Connect, TLS, Server Wait and Transfer
“The API is slow” starts a hunt. Somebody opens the APM dashboard, somebody else greps the load balancer logs, a third person checks whether the database is on fire. An hour later the team has three theories and no measurement. One request, timed phase by phase, answers the first question in under a second: whether the time went to name lookup, the TCP connect, the TLS handshake, the server or the transfer of the body.
That’s the whole proposal. A small command-line tool sends a few requests to one URL, splits each into those five phases, and prints a verdict in plain English: “Server wait is 93% of a warm request. Start with the server.” The phase numbers aren’t new. The verdict is the product, because someone reading a breakdown for the first time can’t tell which number matters, and a few rules can. Nothing here needs a dashboard, an agent or an account. You run it from a terminal while the argument is still going.
The Numbers Already Exist
curl has printed them for years. Five of its write-out variables cover the whole request: time_namelookup, time_connect, time_appconnect (the end of the TLS handshake), time_starttransfer (the first byte of the response) and time_total.
$ curl -s -o /dev/null -w 'dns %{time_namelookup} connect %{time_connect} tls %{time_appconnect} first-byte %{time_starttransfer} total %{time_total}\n' https://api.example.com/orders
dns 0.002 connect 0.010 tls 0.041 first-byte 0.722 total 0.771
Every value counts from the start of the request, so the TLS handshake took 31 ms (0.041 minus 0.010), server wait was 681 ms (0.722 minus 0.041) and the body took 49 ms (0.771 minus 0.722). That subtraction is the first thing people get wrong when they read curl’s output. httpstat, in its Python and Go versions, does the subtraction and draws a timeline, and browser devtools show the same split in the network panel. OpenTelemetry gives the full picture once a service is instrumented, while hey, oha and vegeta measure behavior under load, which is a different job. All of them print numbers and leave the reading to you.
The one phase a client can’t split on its own is server wait, and that’s usually where the time is. The W3C Server-Timing header closes the gap: the server reports its own breakdown in a response header, such as Server-Timing: db;dur=53, cache;desc="miss";dur=2. A client that reads it can say the database took 610 of the 664 ms, with no agent and no vendor. Adding it to a server takes one middleware: start a clock, wrap the database call and the cache lookup, write the header before the response starts. Two cautions apply. Headers go out before the body, so work that happens during streaming can’t be reported. And the header tells every caller, strangers included, how your backend spends its time, so many teams emit it only for requests that carry a debug header or come from an internal network.
What One Run Would Print
A tool built on this idea sends eight requests by default. The first goes out cold, on a fresh connection, and the rest reuse it. It reports the cold request on its own, then the median and the worst value of each phase across the warm ones. It never shows an average, because one four-second outlier wrecks the mean and hides the thing you most want to see. Along the way it reads Server-Timing, the HTTP version, Content-Encoding, body size, the redirect count and the headers CDNs leave behind (Age, cf-cache-status, x-cache). The command name below is only a sketch of the interface.
$ whyslow https://api.example.com/orders --runs 8
GET /orders HTTP/2 200 312 KB no Content-Encoding
cold warm median warm worst
DNS 2 ms 0 ms 0 ms
Connect 8 ms 0 ms 0 ms
TLS 31 ms 0 ms 0 ms
Server wait 681 ms 664 ms 702 ms
Transfer 49 ms 47 ms 58 ms
Total 771 ms 711 ms 760 ms
Server-Timing: db 610 ms, app 41 ms (651 of 664 ms accounted for)
CDN: none detected
Verdict: server wait is 93% of a warm request. Server-Timing puts 610 ms
of it in db, so start with the queries behind /orders.
Also: the 312 KB body is sent uncompressed. Transfer is cheap on this
link and won't be on a slower one.
The verdict comes from a short list of rules, written down so people can argue with them. The thresholds are guesses that need tuning against real services.
- Server wait above about 80% of a warm request means the server is the problem. Name the biggest
Server-Timingentry if there is one, and say so if there isn’t. - A TLS or connect cost on every warm request means the server, or a proxy in front of it, closes the connection after each response. Real clients can’t reuse connections either.
- Transfer that dominates a large, uncompressed body points at compression, or at returning less data.
- Transfer that dominates a small body points at a server that streams slowly, so treat it as server time.
- DNS or connect cost on the cold request only is the price of a first visit, and the network or the resolver owns it.
- Anything else gets “can’t tell”.
The Hard Part Is Knowing What You Measured
Cold and warm are different measurements. The first request after a pause pays for a DNS lookup, a TCP connect and a full TLS handshake. Everything after it benefits from caches, session resumption and an open connection, for reasons that have nothing to do with your server. The two get separate columns, and the tool never blends them.
One sample lies. A request can land on a cold cache, a garbage collection pause or a noisy neighbor, and a single-shot tool will blame the wrong phase with total confidence. Eight runs with a median and a worst case show the typical request and the spread. If the worst is several times the median, the verdict should call the service erratic before it says anything else.
Server wait also contains a network round trip. The request has to reach the server and the first byte has to come back, so a fast server far away looks slow. The TCP connect on the cold request costs about one round trip, which makes it a fair estimate; subtract it from server wait before blaming the server. This is also why running the tool from a laptop on hotel Wi-Fi measures the Wi-Fi. The tool should print where it ran, and the honest routine is two runs, one from a machine near the server and one from where the users sit. The near-server run is mostly the server. The gap between the two is the network.
A CDN in the path changes what you timed. You’re measuring the edge, not the origin, and a cache HIT can look wonderful while every MISS takes two seconds. The tool reads Age, cf-cache-status and x-cache, reports HIT or MISS, and refuses to give a server verdict from a HIT.
HTTP versions don’t time the same way. HTTP/2 over TCP keeps the familiar phases. HTTP/3 runs over QUIC, which folds the transport and crypto handshakes into one exchange, so “connect” and “TLS” can’t be separated cleanly. That’s why 0.1 stops at HTTP/2.
Streaming breaks the usual reading. A server that flushes headers early makes the first byte arrive fast, so server wait looks tiny, and then the body trickles in over three seconds. A naive tool blames the network. A small body with a long transfer means the server is generating the response as it sends it, and the verdict should say that.
Keep the verdict honest. “Can’t tell” is a legitimate answer: high variance, a CDN hit, or a long server wait with no Server-Timing to split it. A tool that always names a culprit teaches people to ignore it.
What Version 0.1 Refuses to Do
Version 0.1 is a single binary that speaks HTTP/1.1 and HTTP/2, prints the phases, parses Server-Timing, detects cache headers, writes the verdict and offers --json for scripts. It refuses to grow a dashboard, to monitor continuously or to generate load. Load testing is a different job with its own traps, which benchmarking under load covers, and mixing it in would make both tools worse.
The verdict is also a handoff. When it names the database, the next tool is a query budget that fails the build when a query scans too much. When it names the transfer, the fix is trimming the payload to the fields clients use. When it names the server and you have no traces, that’s the moment to wire up logs, metrics and tracing, with a specific question to answer. The same order holds inside the application, where profiling before optimizing beats guessing, and where you should pick the metrics that actually matter before building a dashboard for them.
Time one request first. The dashboards can wait.