Continuous Batching Mechanics and Head-of-Line Blocking
Continuous batching solves one head-of-line problem but leaves two deeper mechanisms intact.

Continuous batching closes the original gap in GPU inference scheduling, but the fix covers less ground than most practitioners assume. Static batching treats the GPU as a single unit of work: the whole batch launches together, runs until its longest sequence finishes, and only then admits anything new. If one request runs long, it holds every other slot in that batch hostage, and the GPU sits partly idle behind it the whole time. The waste compounds because requests vary enormously in length. One slow sequence in a batch of thirty-two taxes the entire batch's completion time, even though the other thirty-one finished their work early and have nothing left to do but wait.
Continuous batching changes the unit of scheduling from the batch to the iteration. After every forward pass, the scheduler checks which sequences have completed, frees their KV cache blocks, and pulls waiting requests into the next step. The structural reason this gain is large comes down to how differently prefill and decode consume GPU resources. Prefill is compute-bound: it runs one forward pass across the entire prompt, saturating the GPU's arithmetic units. Mixing these two phases without scheduling at the iteration level wastes both resources at once, since neither phase can use the GPU the way it needs to while the other occupies the loop.
What continuous batching removes is the coarse, batch-level form of head-of-line blocking: a short request sitting idle behind a long one purely because the two arrived inside the same batch window. That's a genuine and well-documented improvement. But it does not follow that continuous batching solves head-of-line blocking as a general problem in LLM serving. It resolves one specific mechanism of the problem, operating at one specific granularity. Two other mechanisms operate beneath it, and they persist in continuous batching deployments running in production today.
The two layers of HOL blocking that continuous batching leaves intact
Continuous batching eliminates head-of-line blocking at the batch level, but it leaves two deeper layers fully intact: prefill-decode interference inside the iteration loop itself, and KV cache memory pressure that can stall request admission.
The first layer concerns resource contention rather than scheduling granularity. Prefill and decode have fundamentally different resource profiles, and forcing both through the same iteration loop creates contention regardless of how finely that loop is scheduled. The second layer concerns capacity rather than scheduling. When GPU high-bandwidth memory (HBM) fills with KV cache blocks, the waiting queue cannot be admitted, regardless of how short the waiting requests are. Clairvoyant documents this failure mode directly: under serial, first-come-first-served backends at high utilization, in environments where continuous batching is infeasible, head-of-line blocking times reach minutes for queries that should return in seconds.
Both layers hide behind a metric that looks fine on its surface. Average tokens-per-second can stay high even as both problems worsen, because averages smooth over exactly the tail behavior these two layers produce. The cluster looks healthy by that measure, but p99 TTFT and queue wait times degrade underneath it. Tail latency and admission-control rejections are what eventually reveal the real picture in production metrics, often well after the problems began.
TTFT spikes from prefill-decode interference inside a single iteration loop
Co-locating prefill and decode inside the same forward pass causes each phase to degrade the other's performance, and the degradation lands unevenly. Decode latency, measured as TTFT for requests already running, absorbs the larger share of the damage. Prefill is compute-intensive and expands the token budget an iteration consumes. When a large prefill batches alongside active decode sequences, it lengthens that iteration, and every decode sequence sharing it experiences a longer step as a direct consequence. Their TTFT grows in proportion to how much prefill work got packed in alongside them.
Disaggregated architectures, where prefill and decode run on separate hardware pools, don't escape this problem so much as transform it into a different shape. The prefill phase in these systems behaves as a non-preemptive, discrete batch processor: once a forward pass begins, the engine locks and cannot accept new inputs until the current batch completes. Staggered Batch Scheduling (SBS) research identifies this locked-batch behavior as a direct source of severe in-engine queuing and parallelization bubbles that degrade TTFT. The request waits for the entire batch ahead of it to finish, no matter how light its own prefill workload is.
Immediate dispatching causes this: it sends a new request to a prefill instance under the assumption that the instance offers continuous service and can absorb work at any moment. The failure compounds further in large-scale deployments that use data-parallel plus expert-parallel architectures, like the ones running DeepSeek-V3 on H800 clusters. An iteration-level scheduler, built for a different problem, was never designed to see the parallelization bubbles that result.
SBS answers this by deliberately buffering requests rather than dispatching them immediately, using the resulting scheduling window to form better execution batches. Those numbers come from a production deployment, not a simulation: they confirm that prefill-decode interference is a measurable tax that shows up in real TTFT distributions at real scale, and buffering with load awareness is what brought that tax down.
KV cache memory pressure as a second queue bottleneck for admission
When GPU HBM fills with KV cache blocks, the continuous batching scheduler loses the ability to admit new requests, no matter how precisely it manages iteration timing. The memory wall takes over as the governing bottleneck, so the scheduling loop can't do anything until capacity frees up again.
PagedAttention, introduced with vLLM at SOSP 2023, addresses a related but distinct problem. Its virtual-memory model removes internal fragmentation by allocating KV blocks on demand rather than pre-reserving a contiguous block of VRAM for each request. That gain matters, because it lets more requests share a fixed pool of memory without wasting space on unused reservations. But PagedAttention does not remove the capacity ceiling itself. It only delays the point at which that ceiling gets hit. Once every physical KV block fills up, new requests queue at admission no matter how short or urgent they are, and paged allocation changes none of that.
FIFO ordering at this layer reproduces the exact head-of-line pattern continuous batching was built to eliminate, just one level down. Clairvoyant documents this as the core failure mode in serial and memory-constrained deployments, with head-of-line blocking times reaching minutes under sustained saturation.
The scheduling response to this problem is shortest-job-first (SJF) admission: you predict how long a request will run before you decide whether to admit it. Clairvoyant implements predictive SJF as a sidecar proxy sitting in front of serial backends such as Ollama and llama.cpp. It predicts response length from 19 lightweight lexical features, using an exported gradient-boosted tree classifier that runs in 0.029 milliseconds per request, and it captures the large majority of the ranking fidelity that a full fine-tuned transformer predictor would offer, at a fraction of the computational cost.
A constraint shapes the entire design: SJF here has to be non-preemptive, because preempting a running job discards whatever KV cache it has already computed. The admission decision has to happen before the job starts. Prediction speed therefore matters as much as prediction accuracy. Curated instruction datasets like Alpaca and CodeAlpaca make poor training sources for a length predictor, because brevity constraints imposed by the model used to generate their output templates reduce long-response examples to a vanishingly small share of the data. Natural conversation logs are the only training signal that holds up for a length predictor meant to run in production, since they capture the genuine spread of short and long responses that real traffic produces.
A second response to the same memory wall works at the hardware layer rather than the scheduling layer: extending KV storage beyond HBM into a three-tier hierarchy. When GPU HBM is exhausted, cold KV blocks can move to CPU DRAM, and from there to NVMe SSD, expanding effective KV capacity without adding another GPU to the cluster. NVMe offloading trades access latency for capacity, so offload bandwidth must keep pace with KV production during prefill, a constraint the next two sections address directly.
Prefill-decode disaggregation as the architectural response to both layers
Separating prefill and decode onto dedicated hardware pools eliminates the interference source behind Layer 1 and gives each phase its own memory pool to manage independently. It also introduces a new binding constraint in its place: the KV cache transfer between prefill and decode nodes.
The premise behind disaggregation follows directly from the resource profiles described earlier. Prefill saturates compute. Decode saturates memory bandwidth. Running both on the same hardware forces that hardware to serve two incompatible resource profiles at once, and no amount of iteration-level scheduling changes the fact that the two phases want different things from the same chip at the same time. If you give prefill and decode dedicated pools, each phase gets sized and scheduled against its actual bottleneck instead of a compromise between the two. DistServe introduced this paradigm, and Sarathi-Serve extended it with chunked prefilling and stall-free batch scheduling, reporting substantial throughput-latency improvements over conventional vLLM deployments.
StreamServe Mooncake builds on the same disaggregated foundation with a KV-cache-centric design, emphasizing prefix-cache-aware scheduling across prefill nodes that prioritizes cache locality by reusing KV data from cached prefixes rather than recomputing it. StreamServe adds two things on top of the basic disaggregated split: metric-aware routing and adaptive speculative decoding. FlowGuard, its router, combines cache reuse, memory utilization, queue depth, and active load signals to route requests across disaggregated compute lanes, supplying a scheduling layer that disaggregated architectures otherwise lack on their own. SpecuStream adjusts speculation depth at runtime based on acceptance-rate gradients, system load, and throughput targets, avoiding the fixed-depth speculation problem that performs poorly the moment load becomes variable.
Those figures show disaggregation works as an architecture at production scale.
Disaggregation's gains come at the price of a new engineering problem: moving the full KV cache for every prompt token from the prefill node to the decode node before decode can begin. For long-context or multi-turn agentic workloads, that KV cache is a large tensor that has to cross a network boundary before any decode token gets produced. The standard CPU-driven RDMA path for that transfer traverses six hops: GPU HBM to a pinned DRAM bounce buffer via cudaMemcpy, then to the NIC, across the fabric, to the remote NIC, into a remote pinned DRAM bounce buffer, and finally into decode GPU HBM. Each hop adds latency and CPU involvement that scales with context length, so the problem grows precisely in the workloads where disaggregation's other benefits matter most.
Disaggregation is not a universal upgrade. On short-context, low-concurrency workloads, the KV transfer overhead can exceed whatever interference savings disaggregation would otherwise provide, which makes the choice workload- and scale-dependent rather than a default best practice. Viability comes down to one question: can the KV transfer path move data fast enough to matter, and the next section takes that up directly.
KV transfer bandwidth and lossless compression at scale
Disaggregation holds up at scale only if the KV transfer path keeps pace with KV production during prefill. To close that gap you need both high-speed interconnect and lossless compression built specifically for AI-native tensor formats, not general-purpose compression tools repurposed for the job.
Prefill generates a large volume of data that the transfer path must keep up with. Every prompt token produces a key and a value vector for every attention head in every layer of the model, and at long context lengths that adds up to a transfer volume that grows linearly with sequence length. GPU-direct RDMA paths, where the NIC reads straight from HBM, remove the bounce buffer hop entirely, but they require careful integration with whatever serving framework sits on top of them, and that integration work is non-trivial.
Compression offers a second lever, but only if it respects what KV tensors actually look like. In BF16 tensors, a pattern first shown for model weights and since confirmed for KV activations, most of the exploitable redundancy is in the exponent field, while the mantissa compresses poorly. Even when ported to run on GPU, these codecs yield little compression benefit against raw floating-point tensor layouts, because they were never designed with that structure in mind.
Lossy compression offers a shortcut around the bandwidth ceiling, but it carries a cost that disqualifies it from the critical path in most deployments. Effective compression for the KV transfer path has to stay lossless, has to run on GPU rather than CPU, has to be format-aware with respect to BF16 and FP8 exponent structure, and has to integrate into the serving framework without introducing a synchronization point that stalls the prefill pipeline it's meant to support.
Storage throughput, once a backend concern handled far from the inference path, now functions as a first-class variable in whether a disaggregated serving architecture meets its latency targets. The three layers trace a single throughline: static batching's coarse blocking gives way to continuous batching's iteration-level fix, which in turn exposes prefill-decode interference and KV memory pressure as the two layers still left standing, and disaggregation answers both at the cost of a transfer problem that only fast, format-aware, lossless compression and genuine GPU-to-GPU bandwidth can resolve.


