Learn how to measure goodput, tune batching and KV cache in vLLM, and decide when another GPU will help, all without breaking your latency SLOs.
Most growing LLM applications eventually face the same problem: they need to handle more users and longer prompts while keeping response times and serving costs acceptable. The usual first lever is batching, which raises throughput by letting more requests share the GPU at once.
A deployment can complete more requests per second while each user waits longer for the first token and the full response, so a better throughput number can hide a slower product. Measure capacity as the traffic you can serve while time to first token and total response time stay inside your latency budget.
This article explains how to establish that budget and measure capacity against production traffic. With those measurements, you can pinpoint what's capping your capacity and whether tuning the server is enough or additional GPUs are needed. The examples use vLLM, but the same method applies to SGLang and other serving engines. The controls just go by different names.
Part of the latency budget is spent before a request ever reaches the inference server, on authentication, quota and billing checks, routing, and network transit. Parasail's guide to reducing LLM latency in production breaks down that request path stage by stage. Here, we'll focus on the inference server and how much traffic it can handle in the time that's left.
Step 1: Define what good latency looks like
Before tuning the server, decide what kind of slowdown would hurt your application.
Consider a coding assistant. A user might tolerate a longer answer if it starts quickly and streams smoothly. Waiting several seconds before anything appears is a different problem from getting the first token immediately and then watching generation crawl.
If you want to improve throughput without making that experience worse, you need different metrics that separate those behaviors. NVIDIA’s NIM benchmarking guide defines the metrics to measure throughput and latency:
- Request throughput (requests/s): How many requests the system completes per second.
- Token throughput (tokens/s): How many output tokens the system generates per second across active requests.
- Time to first token (TTFT): How long a request waits before the first output token arrives.
- Inter-token latency (ITL): The average delay between output tokens after generation begins. ITL is closely related to time per output token (TPOT), and some tools report both.
- End-to-end latency: How long the complete request takes from submission to the final output token.
But how do you track those latency measurements? A p50 median latency is useful, but requests don’t all experience the same conditions. Some arrive when the server is relatively free, while others wait behind existing work or compete with more active requests. The median can hide these slower cases.
That’s why you need to look at tail latency. A p95 TTFT of 800 milliseconds means that 95% of requests receive their first token within 800 milliseconds, while the slowest 5% take longer. Looking at p95 or p99 reveals whether a higher load is creating a small but important group of much slower requests.
Once you’ve defined the latency metrics that determine success for your product, you can set the service-level objectives (SLOs). For a coding assistant, you may want p95 TTFT below 1 second and p99 end-to-end latency below 10 seconds. An offline summarization job may accept much longer latency for better utilization, while real-time products like voice agents need much tighter targets.
Those targets cover the whole request, so split them by stage before you tune the server. When we built sub-second inference for global AI traffic, we started from a 600ms end-to-end p99 and gave 300ms to network and gateway, 250ms to model serving, and 50ms to margin. The serving share is the number your benchmarks have to meet.
These SLOs define what good looks like for your product, and a throughput gain only counts if it stays inside them. The next step is to find how much production-like traffic the deployment can sustain before it crosses those limits.
Step 2: Benchmark LLM inference with production-like traffic
Once the ideal latency SLOs are established, you can move on to assessing the reality. How much traffic can the current deployment handle before it starts breaking them?
The answer depends on the workload. A benchmark with short, uniform prompts may show strong throughput and acceptable latency. Longer prompts, variable output lengths, and bursts of requests can push the same deployment past its latency targets at a lower request rate. Before you trust an optimization, benchmark it against traffic modeled on your own production logs.
Reproduce the production workload
Start with the requests themselves. Prompt length determines how much work happens during prefill, while generated output length affects decode time and how long KV-cache state remains active.
For example, a coding assistant might see everything from a short question about a function to a request containing several files of repository context. The large request spends far longer in prefill and holds much more KV-cache memory, so a benchmark built around the average prompt length misses the requests that put the most strain on the server.
When possible, reproduce the production distribution of input and output lengths, request types, arrival rates, concurrency, and traffic burstiness. If production traces are available, anonymized traces are the closest representation of the real workload. Otherwise, choose a public dataset that resembles the application. Hugging Face Datasets includes conversational, coding, summarization, reasoning, and other datasets that can serve as starting points.
| Application | Dataset | Input Length | Output Length |
|---|---|---|---|
| Chat assistant | Short | Short | |
| Coding assistant | Short | Long | |
| Document Summarization | Long | Short | |
| Reasoning assistant | Short | Medium |
These categories are illustrative rather than fixed properties. Exact sequence lengths depend on the examples selected, the tokenizer, and any filtering or truncation applied during benchmarking. The goal is to choose and sample a dataset whose request shape resembles the application you want to serve.
Increase the load gradually
Once the workload looks realistic, start increasing the pressure on the server.
vllm bench serve controls the request rate, burstiness, and maximum concurrency. Its default request rate is inf , which sends requests as quickly as possible and is useful when you want to push the server toward maximum throughput. For a production capacity test, you want to reproduce the arrival pattern your application sees.
Start below the expected capacity and increase the offered load. At first, more concurrent requests may give the scheduler more opportunities to batch work and increase throughput. Eventually though, requests begin waiting behind other work. Throughput starts to flatten while latency continues to rise.
That’s the transition you’re looking for.
Measure goodput, not just throughput
Goodput is the number of requests per second that complete within your latency SLOs. Raw throughput counts every successfully completed request:
Throughput = completed requests / second
Goodput counts only the requests that complete while satisfying the latency SLOs:
Goodput = completed requests that meet SLOs / second
Throughput tells you how much work the server finishes. Goodput narrows that count to work that finishes while still meeting the application’s latency requirements.
Consider an illustrative sweep of a coding assistant:
| Concurrency | Throughput | Goodput |
|---|---|---|
| 16 | 14 req/s | 14 req/s |
| 24 | 17 req/s | 12 req/s |
At concurrency 24, the server completes more requests overall, but misses more of the latency SLOs. Raw throughput rises from 14 to 17 requests per second, while goodput falls from 14 to 12. The higher-concurrency configuration has more throughput, but less usable capacity for this workload.
This is why you increase load rather than benchmark at a single concurrency level. You want to find the region where goodput stops improving and latency begins to move outside your budget. That establishes a baseline for the optimizations that follow. Batching, memory management, precision, and additional GPUs now have something meaningful to beat.
Step 3: Tune batching and concurrency together
Batching lets the GPU process work from multiple requests together. Higher concurrency gives the scheduler more requests to batch, so aggregate throughput usually rises as concurrent load increases, up to a point.
That point is usually near the engine's maximum batch size. Once the batch is full, throughput levels off and extra requests wait in a queue, which adds latency without adding output. NVIDIA recommends sweeping concurrency from one request to slightly above the engine’s maximum batch size so you can see what happens for your workload.
Track TTFT and ITL at each concurrency level alongside throughput. Prefill processes the input and builds the KV cache, so it drives TTFT. Decode produces output one token at a time, so it drives ITL and how smoothly the response streams. Because the two phases stress the GPU differently, a change can improve one metric while degrading the other.
vLLM V1 uses chunked prefill by default when supported. The scheduler prioritizes active decode requests, then uses the remaining token budget for new prefills. If a prompt is too large to fit, vLLM splits its prefill into smaller chunks. This lets compute-heavy prefill work share a batch with memory-bound decoding, improving GPU utilization without letting long prompts stall token generation.
Tune this behavior with max_num_batched_tokens :
- To reduce ITL: Use a smaller value so fewer prefill tokens compete with active decode requests.
- To reduce TTFT: Use a larger value so vLLM can process more prompt tokens in each iteration.
- To optimize throughput: For smaller models on large GPUs, vLLM recommends testing values above 8,192.
Treat these values as starting points. Sweep the token budget and concurrency using your production request mix, and measure TTFT, ITL, and goodput for each configuration. Stop increasing batching when it produces little additional goodput or pushes latency percentiles outside the budget.
Step 4: Manage KV-cache memory pressure
Increasing concurrency only helps while the GPU has enough memory for the active requests. Each request stores KV-cache state for tokens already processed, so longer contexts, longer outputs, and more simultaneous requests all increase the memory footprint. Eventually, memory can become the limiting resource even when the GPU still has compute capacity available.
vLLM manages KV-cache memory with PagedAttention, which divides the cache into blocks and allocates them as sequences grow rather than reserving one large contiguous region for each request. This reduces fragmentation and allows more requests to remain active concurrently.
However, memory is still finite. When KV-cache space becomes insufficient, vLLM can preempt requests and recompute their state when capacity becomes available. Recomputation keeps the server running, but it also adds work and can increase latency.
If higher concurrency mostly increases KV-cache utilization, preemption, queue length, and tail latency without improving goodput, you’ve reached a memory limit rather than found more useful capacity. At that point, consider reducing the maximum sequence length or concurrency, adjusting max_num_seqs or max_num_batched_tokens , or increasing available memory.
Quantization can raise that limit too. Lower-precision weights take less memory, leaving more room for KV cache, but the speed impact depends on how well your hardware supports the format. Quantizing the KV cache itself reduces that memory directly. vLLM supports several quantization formats, with support and performance depending on the hardware and implementation.
Measure each effect on its own. GPU metrics will show the memory savings right away. Serving speed and quality take more work, so rerun the concurrency sweep on your production hardware and compare outputs against your quality bar. Keep the quantized model if it serves more requests within your SLOs and your quality checks still pass.
Step 5: Scale with replicas or tensor parallelism
If your deployment still can't serve your target traffic within its SLOs after tuning, it's time to add GPUs. How you use them depends on why you ran out of room.
First check model fit and KV-cache headroom. vLLM's scaling guidance starts with a single GPU when the model fits. If the model fits comfortably on one GPU and you need to serve more independent requests, add a replica. With data parallel deployment, vLLM runs separate copies of the model that each process their own batches with their own KV cache, so replicas don't need to communicate during inference for dense models.
If the model is too large for one GPU but fits within one multi-GPU node, vLLM recommends tensor parallelism (TP), which shards each model layer across GPUs. Splitting the weights can also leave more memory per GPU for KV cache.
TP adds communication to the inference path which makes the interconnect part of the performance decision. vLLM specifically notes that on systems without NVLink, pipeline parallelism can be preferable in some configurations because it reduces communication overhead relative to TP.
For a coding assistant, suppose the model fits on one GPU and the current instance reaches its latency limit during a traffic spike. A second replica may be more useful than splitting every request across two GPUs because it creates another independent serving queue. If the same model barely fits and KV-cache pressure is already causing preemption, sharding may solve a different problem by freeing memory per GPU.
Match the serving architecture to the workload
Once you know your workload's latency budget and where capacity runs out, you can choose infrastructure based on your own numbers.
Start by finding the highest goodput your current configuration can sustain within your SLOs. Then identify whether compute, KV-cache memory, communication, or traffic variability is limiting it. That tells you whether the next dollar belongs in a configuration change, another replica, model parallelism, different hardware, or a different serving model.
Traffic shape then decides the serving model. A steady workload keeps dedicated GPUs busy, while traffic with sharp peaks and long quiet periods pays an idle GPU tax between spikes. Async work can trade latency for higher utilization, and interactive traffic needs enough ready capacity to protect tail latency.
Parasail's managed inference framework compares serverless, dedicated capacity, Elastic Endpoints, and batch against your latency requirements, traffic shape, and configuration needs.