Diagnose LLM latency before you tune it. Split TTFT, TPOT and end-to-end latency, map them to eight stages, and find which one owns your slowdown.
A request that waited 400ms in a queue and a request that spent 400ms in prefill arrive at the client as the same number, but different parts of the stack own them. Queue time belongs to capacity and routing, while prefill time belongs to the engine config. Pulling a lever before you know where the slowdown is happening is a guess, and a guess isn’t a strong enough foundation to build a product on.
Latency is measured in three numbers:
- Time to first token (TTFT) is the pause before anything appears
- Time per output token (TPOT) is the pace of streaming after that
- End-to-end latency is the whole request
The same optimization can move TTFT and TPOT in opposite directions, and which number matters depends on the product. A chat user, or someone watching a model reason, sees TTFT and then TPOT. An agent that can't start its next step until the answer finishes cares about end-to-end latency.
This guide splits a slowdown into those three and maps them onto the eight stages a production request crosses. Once you know what’s slowing you down, our throughput guide covers how to tune the engine with batching, KV cache and GPU allocation.
p50, p90, and p99 tell you what kind of problem you have
The percentile where a slowdown starts tells you roughly what share of requests it touches. When p50 rises, at least half of requests got slower, and the cause is something most of them share, such as a longer network path, a replica running saturated all day, prompts that grew or a config change.
When p50 holds and p90 rises, at least one request in ten is affected, which usually points to one region, replica, or class of long prompts. When only p99 rises, something rare is hitting around one request in a hundred.
The familiar rare events split by metric. Head-of-line blocking, where ordinary requests queue behind rare long ones, and cold starts during scale-out add to TTFT and never touch TPOT. Preemption mid-generation and speculative decoding on requests with low draft acceptance add to TPOT.
Two less obvious causes fill the tail too.
- Hardware variance. Nodes with the same GPU are not interchangeable. We've measured meaningfully different latency between hosts with identical GPUs and different CPUs, an Intel Xeon 6767P against a 6972P. Because a bad node carries only its share of traffic, the fleet median looks healthy while that node's requests fill the tail on both TTFT and TPOT, at p99 in a large fleet and as high as p90 in a small one.
- Prefix cache misses for agents. An agent sends its whole conversation with every call, so a prefix cache hit is what keeps that growing context from being prefilled from scratch. A miss, from landing on a replica that doesn't hold the prefix or from eviction, forces the engine to prefill the entire conversation again. Misses are a small share of calls, which puts them in the tail, and each one adds more TTFT the longer the conversation has run.
Diagnose across all three percentiles, and choose the one you hold yourself to separately. Holding p99 means keeping spare capacity for the worst 1% of cases, and that capacity sits idle the rest of the time, while a p95 target needs less of it and leaves the median almost unchanged. A voice agent, where every slow reply is a dead pause, may need p99. A batch job or an internal tool usually doesn't.
Half the latency budget sits above the engine
When we rebuilt our request path for sub-second inference on global traffic, we worked backward from a 600ms end-to-end p99 on a latency-sensitive workload class where the tail was worth paying for. The allocation came out at 250ms for the model-serving path and 300ms for network and gateway, with 50ms of margin.
The network share covers two logical hops in each direction. The first runs from the client to the gateway, where the API key, quota and billing policy are evaluated. The second runs from the gateway to the replica serving the request. Both round trips are paid before the model starts work, and multi-region LLM deployment with gateway architecture covers which responsibilities belong at the gateway.
Geography sets the floor under that share. A US-East to APAC round trip adds roughly 180 to 220ms at p50 before any model work starts, and no serving flag reaches it. That floor decides how much of the budget is left for tuning at all, so a team serving one region and a team serving three are working on different problems even at the same target.
Routing is one of the most important layers above the engine for TTFT. For each request, the router picks a replica, and that pick decides whether the request hits the prefix cache, whether it waits in a queue, and how large a batch it joins. Each of those shows up as engine time, so bad routing looks like slow prefill, long queues and high TPOT on replicas that each seem fine.
Prefix affinity and load balancing pull against each other. Sending each conversation back to the replica that holds its cache can overload that replica, while spreading requests evenly causes the agent cache misses described above. Weighing both on every request takes a more complex routing layer than round-robin, and ours holds prefix cache hit rates at 85% on most of our major endpoints.
Where an LLM inference request spends its time
The latency budget splits across eight stages. The first six are crossed before the first token arrives, so a delay in any of them shows up as TTFT. Decode sets TPOT, and the response path can make steady decode look uneven by the time tokens reach the client. Together they add up to end-to-end latency.
| Stage | What happens | Where you see it | What drives it |
|---|---|---|---|
| Client to edge | DNS, TLS, transport | Client-side timing | Regional placement, edge termination, connection reuse |
| Edge to regional gateway | Route selection across regions | Edge and gateway timing | Health-based routing, measured route quality |
| Gateway | Authorization, quota, policy, logging | Gateway spans | Gateway runtime, decision caching |
| Cluster ingress to host | Internal networking to the serving replica | Internal spans | Taking the data path off the ingress stack |
| Admission and queueing | Waiting for a scheduling slot |
| Replicas, concurrency limits, capacity headroom |
| Prefill | Processing the input prompt |
| Prefix caching (TTFT, and the TTFT tail for agents), chunked prefill (later first token on long prompts), parallelism strategy (p50), quantization |
| Decode | Generating output tokens |
| Chunked prefill (TPOT p99), batch size (TPOT), speculative decoding (faster average TPOT, wider TPOT p99), prefix caching (less prefill competing with decode), parallelism strategy (p50), quantization |
| Serialization and response | Returning tokens to the client | End-to-end time less engine time | Streaming, placement |
How to diagnose a latency bottleneck in 5 steps
A blank pause before streaming can come from prefill, queueing, admission or the network path. Uneven streaming can come from decode, from prefill interrupting decode on a shared batch, or from connection buffering that never touched a GPU. Run the steps in order and each one rules out a stage the one before it could not see.
The metric names here are vLLM's. We run both vLLM and SGLang in production, and vLLM is where most open-model deployments start. The stages are engine-agnostic, and SGLang and TensorRT-LLM expose equivalents under their own names. vLLM's own names moved between its V0 and V1 engines, and counters append a _total suffix in the exposed text format, so check them against the release you run before you build panels on them.
1. Split end-to-end latency into TTFT, TPOT and output length
End-to-end latency is roughly TTFT plus TPOT for every output token after the first, so it grows with output length while TPOT does not. Read TTFT, inter-token latency and output token counts over the same window and see which one moved. A rise in TTFT sends you through steps 2 to 5. A rise in TPOT with a steady TTFT puts the delay in decode or on the path that streams tokens back, which steps 2, 4 and 5 separate. If end-to-end latency climbed while both stayed flat, responses got longer and no stage owns the delay.
2. Take timestamps at the client, the gateway, and the engine
Measure TTFT at the client and read vllm:time_to_first_token_seconds at the same time. vLLM's metrics design calculates that metric relative to the frontend arrival_time , which as of September 2026 starts when tokenization begins, so DNS, TLS, transport, edge routing and authorization all sit outside it. The gap between the two series is your network and gateway segment, and no engine dashboard shows it. Gateway spans split that gap further, into the hops before the gateway and the work inside it.
Do the same for streaming. When the client sees gaps between tokens that vllm:inter_token_latency_seconds doesn't, a proxy or load balancer between the engine and the user is buffering tokens and releasing them in bursts.
3. Compare waiting requests against running requests
vllm:num_requests_waiting above zero while vllm:num_requests_running sits at the batch ceiling means requests are queued for a scheduling slot rather than waiting on computation. vllm:request_queue_time_seconds gives you the delay itself rather than the queue depth implying it. Queueing delay lands on time to first token, and no prefill flag shortens a queue.
4. Read KV cache occupancy and the preemption counter
vllm:kv_cache_usage_perc near its ceiling caps how many requests run concurrently, which produces the queue from step three. A climbing vllm:num_preemptions_total means KV cache pressure is forcing the engine to evict running requests. A request evicted before its first token adds to TTFT, and one evicted mid-generation stops streaming until it resumes, which adds to TPOT. How much each eviction costs depends on the setup. By default the evicted request's KV cache is recomputed, and with CPU offload, which we run in production, it is reloaded from CPU memory instead.
5. Separate prefill from decode, then segment
vllm:request_prefill_time_seconds measures the interval from scheduling to the first token, and vllm:request_decode_time_seconds measures the interval across successive token outputs. Together they separate a prefill problem from a decode problem once everything upstream is accounted for, and prefill vs. decode explains why the two phases behave differently.
Read decode time against concurrency, since decode slows as the batch grows. Then break the distribution down by the variables that actually differ: source geography, destination region, connection reuse, replica health, prefix cache hit rate, and input and output token distributions. A single global percentile hides a slow geography behind a large volume of nearby traffic, and a p99 driven by the longest tenth of prompts is a different problem from one driven by concurrency or by prefix cache misses.
Cold start and failover produce a tail diagnostic runs don’t show
Both events are rare in steady state, so the steps above won't catch them, and both land entirely on the tail. The burst, termination and scale-out tests in the next section are how you see them.
Cold start: Model load time is invisible in steady state and dominates the tail during scale-out and after a node failure. Our work on faster model loading with fastsafetensors and io_uring cut load time by 3 to 5x.
Failure: A single node failure produces three separate tail events. Detection takes time. Surviving replicas absorb a sudden queue. The replacement instance loads the model and warms before it serves predictably. Redundant replicas are an availability answer to a latency question.
How to prove a latency change worked
A change is proven only when the result holds under production traffic. We validate in three stages:
- Offline benchmark. Run a workload shaped like production through steady load, a burst, a GPU termination and a scale-out event, since a steady-state test never produces the last three. On the termination, measure time to stop routing to the failed node, p95 and p99 during redistribution, queue depth on survivors, and time to warm performance on the replacement. This stage tells you whether the change works under controlled conditions before any real traffic touches it.
- Mirroring. Clone the production deployment onto one or two replicas, apply the change there, and mirror live production traffic onto them. Real traffic carries the prompt mix, burst timing and geography a synthetic workload only approximates, so this stage tells you whether the offline result holds against it.
- Rollout. Roll the change out across the fleet, keep the same per-stage measurements running, and compare them against the baseline from before the change. Keep the previous configuration ready to restore if the target metric or error rate moves the wrong way.
At every stage, record client-to-edge, gateway processing, gateway-to-engine, queue and model time on every request, keeping model execution time and queue time separate, and report distributions by geography rather than one aggregate.
A change succeeds only if it improves the metric it targeted at the percentile you hold yourself to, without regressing the other two, raising errors or making failover less predictable. Throughput measures work the system completes and latency measures time one request spends inside it, so a setting that raises tokens per second while lengthening queues has bought throughput with the tail budget. That's not a latency improvement, however good the headline number looks.
If you want your own latency diagnosed against your token profile, traffic shape and geography, talk to a Parasail engineer.
FAQ
What is the fastest way to reduce LLM latency in production?
Measure before you tune. Splitting end-to-end latency into TTFT, TPOT and output length tells you which part of the request got slower, and comparing client timings with the engine's tells you whether the delay sits upstream of it. Both take an afternoon, and every lever after them is cheap once you know which stage owns the delay.
Why is my client-measured TTFT higher than vllm:time_to_first_token_seconds?
The engine metric starts counting when the frontend receives the request. DNS lookup, the TLS handshake, transport, edge routing, gateway authorization and quota checks all happen before that point. Subtract the engine figure from the client figure and the remainder is your network and gateway time, which for globally distributed traffic is often the larger share.
What are the most common causes of high TTFT in vLLM?
Queueing on a saturated server, KV cache pressure and preemption, long inputs including expanded RAG context, prefix cache misses on traffic that should hit, and everything upstream of the engine. The first two and the last are the most frequent, and a prefill flag fixes none of them.
What causes high TPOT in vLLM?
Large decode batches, long prompts prefilling on the same replica, preemption under KV cache pressure, and speculative decoding on requests with low draft acceptance. Response buffering between the engine and the client produces the same symptom without touching the GPU, so compare client-side token gaps with the engine's inter-token latency before changing any engine setting.
Does Kubernetes add latency to LLM inference?
It can when it sits in the data path. Each ingress, load balancer and service-mesh hop adds a worst case that compounds into p99. Keeping Kubernetes for scheduling and scaling while latency-sensitive requests travel a direct path removes those hops and keeps the orchestration.
How do I monitor LLM inference latency?
Scrape the vLLM /metrics endpoint and build five panels: TTFT quantiles, inter-token latency quantiles, waiting against running requests, KV cache usage, and the preemption counter. Add client-side TTFT and token timing beside them, because no engine metric can see the network, the gateway or the response path.
Does prefix caching help if every prompt is unique?
No. Prefix caching reuses computation for tokens that requests share, so traffic with no shared prefix gets nothing from it. Compare hits against queries before investing, since prompts that look unique often share a system prompt or scaffold.