KV Cache Sizing for Concurrent Request Budgets

How to predict GPU memory limits before your inference server crashes from hidden cache overhead.

Senior Writer · · 11 min read
Cover illustration for “KV Cache Sizing for Concurrent Request Budgets”
KV Cache Systems · September 30, 2026 · 11 min read · 2,549 words

GPU compute utilization is far below saturation while inference servers return CUDA OOM errors, because the constraint is memory. Model weights load fine in testing, so the assumption is that the hard part is done. It isn't. The KV cache grows with every concurrent request instead of getting paid once at load time, and it's the variable nobody sized for.

A transformer has to attend to every prior token to generate the next one, and rather than recompute that attention from scratch at each step, the model stores the key and value tensors from every previous token and reuses them. That reuse is what makes autoregressive decoding fast. It also means memory holds a growing ledger of every token in every active conversation, and that ledger scales along three axes at once: context length, batch size, and concurrent user count. None of those move in isolation in a production system. They compound.

The scale of the problem is easy to underestimate until you run the numbers once. A single Llama 3 70B request at 128K tokens of context needs roughly 42 GB of GPU memory just to hold its KV cache, on top of around 140 GB for the FP16 weights themselves KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron…. An 80 GB card can't even hold the weights and that one request together, let alone serve a second user KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron…. That's the whole problem in one example: the cache isn't a rounding error next to the model, it can dwarf the model.

The formula behind it makes the sizing predictable rather than mysterious. It's arithmetic, not guesswork, and once you can derive your own per-request footprint, sizing a fleet for a concurrency target becomes a planning exercise rather than a trial-and-error scramble.

The exact formula that governs KV cache footprint

The leading 2 is there because the cache holds both the key tensor and the value tensor for every token.

Layer count and KV head count (the head count that survives after grouped-query attention reduction, not the raw query head count) do almost all the work in determining the final number. Total parameter count barely matters by comparison. A 70B model and a much smaller model running the same grouped-query attention configuration can land in very different places on KV footprint, because footprint tracks architecture choices, not headline size.

Running the Llama 3.1 family through the formula at BF16 makes this concrete. The 8B model, with 32 layers, 8 KV heads, and a 128 head dimension, costs about 0.131 MB per token KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron…. The 70B model, at 80 layers with the same 8 KV heads and 128 head dimension, costs about 0.327 MB per token KV Cache Optimization for LLMs 2026: Engineering Guide KV Cache Memory: The Real Cost of Long-Context Inference | IntuitionL… KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron…. The 405B model, at 126 layers, comes to roughly 0.516 MB per token What Databases Knew All Along About LLM Serving. Notice that going from 8B to 405B (32 layers versus 126 layers), per-token cost moves from about 0.131 MB to about 0.516 MB, because KV heads stayed fixed at 8 the whole way KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron… KV Cache Optimization for LLMs 2026: Engineering Guide KV Cache Memory: The Real Cost of Long-Context Inference | IntuitionL… What Databases Knew All Along About LLM Serving.

That fixed head count is grouped-query attention doing its job.

Models that use sliding window attention, some of Mistral's layers among them, cap the per-layer KV cache at the window size rather than letting it grow with the full context, so the plain formula needs adjusting for those architectures. And DeepSeek's multi-head latent attention design doesn't cache full key and value tensors at all: it compresses them into a latent vector first, and DeepSeek's own paper reports a 93.3% KV cache reduction against a comparable dense architecture KV Cache Memory: The Real Cost of Long-Context Inference | IntuitionL…. The formula's structure, not just its inputs, changes for MLA models KV Cache Memory: The Real Cost of Long-Context Inference | IntuitionL…. KV_bytes = 2 × L × H_kv × D × S × B × bytes_per_element. bytes_per_element = 2 for BF16/FP16, 1 for FP8, 0.5 for FP4. Llama 3 uses GQA with 8 KV heads instead of 64 query heads, and that 8× reduction in KV heads cuts KV cache size 8× versus standard MHA at the same scale, Spheron reports, though this ratio is Llama 3-specific rather than a universal GQA property, as Gemma 2, also GQA, achieves only a 2× reduction.

Translating per-token cost into a concurrency budget

At 4,096 tokens, one user costs about 1.3 GB and eight users cost about 10.7 GB What Databases Knew All Along About LLM Serving. At 16,384 tokens, one user is about 5.4 GB and eight users climb to roughly 43 GB. At 32,768 tokens, one user runs about 10.7 GB and eight users reach around 86 GB What Databases Knew All Along About LLM Serving. Push to 131,072 tokens and one user alone needs about 42.9 GB, with eight users landing near 343 GB⟧c25⟧. Somewhere around the 32K mark with eight concurrent users, the KV cache by itself has already passed the 140 GB the model's weights occupy KV Cache Memory: The Real Cost of Long-Context Inference | IntuitionL…. The cache has become bigger than the model it's serving.

Batch size does the same thing from a different angle.

Concurrency itself needs a precise definition rather than a guess. If requests take 20 seconds each to generate and new ones arrive once a second, roughly 20 requests are in flight at any given moment, since steady-state concurrency is arrival rate multiplied by mean response time. Double the generation time at the same arrival rate and concurrency doubles too, with no change in traffic. Overload makes this worse in a way that's easy to miss during planning: clients that retry after a timeout add more in-flight requests on top of whatever was already there, pushing memory pressure past the ceiling the original estimate assumed.

Long context eventually flips the entire optimization problem. Past that point the system isn't compute-bound anymore in any meaningful sense; it's memory-bound, full stop KV Cache Optimization for LLMs 2026: Engineering Guide. Size for peak concurrency, not average, work out the KV budget first, and then check that what's left of VRAM actually fits the weights. If it doesn't, that shortfall is the real hardware minimum, or the signal that quantization isn't optional. The core multiplication is per-token KV cost × context length = per-request footprint, and per-request footprint × concurrent batch size = total KV budget required. A worked concurrency table for Llama 3.1 70B at BF16 is provided, from Spheron.

Diagram: KV Cache Dwarfs Model Weights at Scale. Visualizes: Show how KV cache memory for Llama 3.1 70B (BF16, 8 concurrent users) grows across four context lengths, eventually eclipsing the fixed 140 GB weight footprint.

Attention architecture choice and its effect on the formula's output

Before any allocator trick or quantization scheme enters the picture, the attention architecture itself has already decided a large chunk of the KV budget. Four variants dominate current practice, and they produce materially different numbers out of the same formula: multi-head attention, multi-query attention, grouped-query attention, and multi-head latent attention KV Cache Optimization: Memory Efficiency for Production LLMs | Introl….

Multi-head attention gives every query head its own KV head, so KV head count equals query head count and the footprint is at its largest possible value. Multi-query attention goes to the opposite extreme: one shared KV head across every query head, which minimizes cache size at the cost of attention quality on some tasks. Grouped-query attention clusters query heads into groups that share KV heads, and it's the standard for large open models across 2024 through 2026 KV Cache Optimization: Memory Efficiency for Production LLMs | Introl…. The reduction it delivers isn't fixed, though: it depends entirely on the grouping a given model chose, not on grouped-query attention as a category KV Cache Optimization: Memory Efficiency for Production LLMs | Introl….

DeepSeek's multi-head latent attention caches a compressed latent vector rather than full K and V tensors, giving a 93.3% reduction in KV cache versus dense architecture per DeepSeek's own paper, and changes the formula's structure entirely. Hybrid designs are also showing up in production traffic now: Jamba 52B mixes transformer, Mamba, and mixture-of-experts layers, and certain Gemma sliding-window variants carry categorically lower KV cost than pure dense attention.

Anyone still choosing a model has, in architecture, the highest-leverage lever available for the KV budget. Anyone locked into a model already has to look elsewhere: allocation, caching, and quantization take over. In MHA, all query heads get their own KV heads, so KV head count equals query head count, producing the largest footprint, and a comparable MHA model at 70B scale would need roughly 8× more KV cache per token than Llama 3.1 70B's GQA config, Spheron reports.

Diagram: How Attention Architecture Shapes KV Footprint. Visualizes: Rank the four attention architectures by their relative KV cache footprint, from largest to smallest: Multi-Head Attention (MHA, baseline — KV heads equal query heads, ~8× larger…

PagedAttention: reclaiming the 60–80% of KV memory that naive allocation wastes

A request that generates 200 tokens against a 4,096-token reservation leaves the remainder parked and unusable. Measured across real workloads, that pattern of fragmentation and over-reservation wastes 60 to 80% of allocated KV memory KV Cache Optimization for LLMs 2026: Engineering Guide.

PagedAttention fixes this by breaking the cache into fixed-size blocks, typically sized at 16 or 32 tokens, and allocating them on demand as generation proceeds rather than reserving everything up front. Blocks get freed the instant a request finishes, so nothing sits idle waiting for a worst-case sequence length that never arrives. The idea borrows directly from operating-system virtual memory: a request sees its cache as one continuous span, while the allocator quietly maps that span onto physical blocks scattered wherever they fit. Measured outcomes back the approach up.

By 2026, this isn't an optional upgrade anymore, it's table stakes KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron…. Every serious production inference stack ships it or an equivalent by default: vLLM and TensorRT-LLM both run PagedAttention, and SGLang ships RadixAttention as its default instead, which solves the same allocation problem through a different data structure. The live question for anyone deploying today is what to layer on top of it.

For teams tuning vLLM directly, --gpu-memory-utilization, commonly set around 0.90, controls what fraction of VRAM left over after loading weights gets dedicated to the KV pool. --max-model-len caps the context window and, by extension, the maximum size of the KV block pool. It should be set to the workload's actual maximum, not the model's theoretical ceiling, since every unused token of headroom there is memory not available to another request. And --kv-cache-dtype sets the quantization format at allocation time, tying allocation directly into the dtype lever covered further down.

Newer eviction techniques stack on top of paging rather than replacing it. PagedEviction removes low-importance blocks at the block level without touching CUDA kernels, and entropy-guided caching goes a step further by allocating cache budget according to each layer's attention entropy, giving more room to layers with broad attention patterns and less to layers that stay narrowly focused. vLLM's own benchmark, cited in Introl and TDS, shows measured outcomes with waste dropping to under 4% and throughput rising 2–4× over naive allocation.

Prefix caching: amortizing shared context across the concurrent population

A large share of production traffic shares context across requests: a long system prompt repeated on every call, a reference document sitting at the top of a RAG pipeline, or the accumulating history of a multi-turn conversation. Without any mechanism to exploit that overlap, every one of those requests recomputes the KV state for the shared prefix from scratch during prefill, and that cost scales with every concurrent user hitting the same shared text.

Prefix caching stores the KV state for that shared prefix the first time it's computed, so any later request starting with the same tokens gets a memory read instead of a fresh computation. The savings scale with how much gets shared: two calls sharing a 200K-token prefix that differ only in 5K tokens of user-specific content turn 200K tokens of attention compute into a lookup.

Field numbers back the concept up without overselling it. Production deployments tracking Prometheus metrics, cache usage percentage, prefix hit rate, eviction rate, effective throughput, have reported warm-cache prefix hit rates around 87% backend.ai.

The gain is conditional, though, and that condition matters more than the headline numbers suggest. Reuse only pays off when requests genuinely share prefixes. Traffic without that structure keeps hit rates low, and the cache ends up spending effort on lookups that mostly miss rather than delivering savings, so hit rate needs to be measured before anyone counts on this as free headroom. RAG pipelines carry their own version of this trap: dumping every retrieved search result into the prompt inflates both KV cache consumption and prefill time, which makes controlling the number and size of retrieved chunks a real design decision rather than a detail to skip over. SGLang's RadixAttention organizes cached prefixes in a radix tree so overlapping prompts share state automatically, and the SGLang team reports up to 6.4× higher throughput on workloads with heavy prefix reuse, TDS reports, a figure that comes from vendor reporting on prefix-heavy workloads. The 2–3× speedup on agentic workloads from prefix caching is consistent enough that it ships as a default in vLLM, SGLang, and TRT-LLM, the prompt20 blog reports KV Cache: The Complete Guide.

Quantization: how dtype choice multiplies or divides the formula's output

Cutting precision in half (for example 2 bytes for BF16/FP16 versus 1 byte for FP8 or 0.5 bytes for FP4) cuts the entire KV budget in half, which at a fixed memory ceiling doubles how many concurrent requests the fleet can hold.

Format choice, though, isn't purely a quality dial teams get to set freely, it's constrained by what the hardware kernels actually support for a given architecture. FP8 works cleanly for Llama-3.1-405B and DeepSeek V3.2. Qwen3-VL-235B can't use it, because its vision encoder dimensions aren't compatible with FP8 block quantization kernels, so it needs BF16. Selecting a dtype, in other words, means checking architecture compatibility first and quality preference second.

Research beyond straightforward format quantization keeps pushing the compression ceiling higher. KIVI, from 2024, runs tuning-free asymmetric 2-bit KV quantization and shows up across multiple 2025 and 2026 survey papers as a workable path to aggressive compression without any fine-tuning step ⚙️ LLM Inference in 2026: Sizing GPUs, KV Cache and Load. KVComp, from 2025, takes a different approach entirely: an LLM-aware lossy compression framework that adapts to the model's internal structure rather than applying one uniform format across the board. And GVote, slated for ICLR 2026, does away with manual budget specification altogether, adaptively compressing the cache instead of forcing every workload into the same fixed-size box, a problem researchers have taken to calling the Procrustes' bed problem. FP8 KV quantization reduces KV memory footprint approximately 50% relative to BF16 for GQA-based models, and vendor documentation reports FP8 KV cache enables 2 to 3× larger batch sizes on H100 GPUs, IntuitionLabs reports arxiv.org. KVquant (NeurIPS 2024) targets enabling extremely long context, framed as a path toward 10M context length, through KV cache quantization, per arxiv:2401.18079.

Sources

  1. KV Cache Memory: The Real Cost of Long-Context Inference | IntuitionLabs
  2. KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron Blog
  3. ⚙️ LLM Inference in 2026: Sizing GPUs, KV Cache and Load
  4. KV Cache Optimization for LLMs 2026: Engineering Guide
  5. KV Cache Optimization: Memory Efficiency for Production LLMs | Introl Blog
  6. What Databases Knew All Along About LLM Serving
  7. KV Cache: The Complete Guide
Filed underKV Cache Systems

More in KV Cache Systems