Chunked Prefill and Its Effect on Time to First Token

Chunked prefill eliminates scheduling bottlenecks that hurt response time.

Staff Writer · · 11 min read
Cover illustration for “Chunked Prefill and Its Effect on Time to First Token”
Inference Serving · October 6, 2026 · 11 min read · 2,392 words

Large language model inference is really two workloads wearing one name, and the industry's failure to separate them has produced serving systems that struggle with both. Prefill, the phase that processes an entire prompt and produces the first output token, is compute-bound: it runs as matrix-matrix multiplication across every prompt token at once, and it saturates a GPU's arithmetic units. Decode, the phase that then generates each subsequent token one at a time, is memory-bound: every step has to read the accumulated key-value cache, and its bottleneck is bandwidth.

Time-to-first-token, the metric users actually feel as "how long until something happens," is set almost entirely by how long prefill takes. It has little to do with how fast tokens stream out afterward. That second quantity, inter-token latency, is a decode-phase measurement, and it responds to a different set of optimizations than TTFT does. Treating the two as one number to optimize is where a lot of serving systems go wrong before they even get to scheduling.

The link between the two phases is the KV cache: a store of attention keys and values that prefill builds and that every decode step reads from, for the rest of the request's life. The cache grows with context length, and that growth is not gentle. Attention computation during prefill scales quadratically with token count, and the memory required for full attention scales the same way, so a long input is not a small multiple of a short one in compute terms, but an entirely different order of magnitude. A single long prefill, in other words, is a real piece of work on the GPU, not a rounding error next to decode.

Naive scheduling and head-of-line blocking from long prefills

Once prefill and decode are understood as distinct computational regimes, the failure mode of naive scheduling becomes obvious. Without chunked prefill, a serving system that receives a long prompt runs its entire prefill as one uninterrupted block of GPU computation. Every other request already being decoded, token by token, in that same batch has to wait. The GPU is a shared resource, and a long prefill claims all of it until it finishes, so the scheduler ends up enforcing head-of-line blocking: one request at the front of the queue stalls everything behind it.

The cost does not land only on the long-prompt request's own TTFT. Every request already mid-decode suffers a latency spike too, because its next token cannot be produced until the GPU is free again, and that produces inter-token latency jitter across the whole batch rather than a clean, predictable cadence. Research describing this problem at the accelerator level makes the mechanism explicit: a long prefill can occupy the hardware for an extended stretch, delaying decodes and shorter requests queued behind it, and introducing exactly the kind of jitter that makes a serving system feel unreliable even when its average throughput numbers look fine.

The obvious fixes do not resolve the problem so much as relocate it. A scheduler that prioritizes prefill gets better throughput but introduces the stalls just described, because it is willing to freeze decoding whenever a prefill request shows up. A scheduler that prioritizes decode avoids those stalls but leaves the GPU underused, because it is unwilling to run prefill at full strength when decode work is light. Agrawal et al., in a 2025 paper published through ACM SIGOPS, frame this as the structural tradeoff that scheduling policy alone cannot escape: the binary choice between prefill-first and decode-first is itself the limitation, and no amount of tuning priority rules resolves it, because the unit of scheduling, the entire prefill job, is simply too coarse. The longer the prompt, the worse this gets: a very long prompt can stall decoding for multiple seconds in naive systems, and as applications keep pushing toward longer context windows, that stall only grows unless the scheduler can operate at a finer grain than "run the whole prefill, then come up for air."

What chunked prefill does mechanically, and where the idea came from

Chunked prefill resolves the problem by changing the unit of scheduling itself. The technique was introduced in the Sarathi-Serve paper by Agrawal and colleagues, published at the 18th USENIX Symposium on Operating Systems Design and Implementation in 2024. Instead of treating a prompt of length L as one monolithic prefill job, the scheduler splits it into ⌈L/S⌉ contiguous chunks of size S and processes them one at a time, interleaving each chunk with a decode step for every request currently active in the batch.

The scheduling discipline that makes this work is called decode-maximal batching: on each iteration, fuse one prefill chunk together with as many decode steps as the batch can supply, so that the GPU runs one large combined matrix multiplication rather than alternating between separate prefill and decode passes. Mechanically, this matters because prefill's compute-bound matmul keeps the GPU's arithmetic pipeline full, while decode's memory reads happen underneath that same computation instead of leaving the pipeline idle while waiting for a dedicated turn. Fusing the two means neither phase sits around waiting on the other, and the GPU spends less time idle overall.

The practical enforcement mechanism is a token budget. Each scheduling iteration gets a fixed budget of tokens, sized to keep the linear layers running near their hardware ceiling. Every decode request consumes exactly one token from that budget. A prefill request consumes tokens equal to its remaining prompt length, unless that would exceed the budget, in which case it is chunked down to fill whatever budget remains. This is the exact mechanism that eliminates head-of-line blocking: no single iteration can run longer than the budget allows, so no prefill chunk, however long the underlying prompt, can freeze decoding for an unbounded stretch. The fix is not a smarter priority rule. The fix is capping the size of the unit that gets scheduled.

The idea has since moved well past its original paper. Chunked prefill is now the default iteration granularity in Sarathi-Serve, in vLLM, and in RServe. It has gone from a research proposal to the standard behavior of the serving frameworks much of the industry already runs.

The chunk size dial and the tradeoff it controls

Diagram: The Chunk Size Dial: TTFT vs. Throughput. Visualizes: Visualize the single tradeoff controlled by chunk size S in chunked prefill scheduling.

Chunk size S is the single dial that determines where a given deployment sits on the tradeoff curve between TTFT and throughput, and the two directions of that dial behave differently. Shrinking S lowers TTFT and reduces inter-token latency jitter, because each scheduling iteration does less work before decode gets its turn again. But smaller chunks mean more scheduling iterations overall, and each iteration carries its own kernel launch overhead, so throughput takes a hit as S shrinks. Growing S improves throughput by amortizing that overhead across more tokens per iteration, but it reintroduces exactly the phase interference chunked prefill was built to remove, because a larger chunk takes longer to run and holds decode requests waiting longer within that same iteration.

Dense models tend to do well with chunk sizes in the 256 to 512 token range, but that range is not universal. Deployments with tight time-between-token service objectives, under roughly 20 milliseconds, often need chunks smaller than that range, because at 256 to 512 tokens a memory-bound decode step can still end up waiting on a compute-bound prefill chunk within the same iteration, and only a smaller S removes that wait. This means throughput and latency targets are not something a deployment sets once at configuration time. They are a scheduling problem that has to be rebalanced iteration by iteration, especially where request lengths vary widely within the same batch.

It would be a mistake to read chunked prefill purely as "a little throughput given up for a lot of latency gained." At high load, piggybacked decoding, the regime where decode steps ride along on every chunked prefill iteration, delivers a large multiple of decode throughput gains for LLaMA-13B on an A6000 GPU, with a modest but meaningful boost to end-to-end throughput in the same configuration. Chunked prefill often improves throughput rather than merely taxing it, because it eliminates the GPU idle time that decode-only stalls would otherwise create. The throughput cost appears mainly at the margins of chunk-size tuning, rather than as a fixed tax the technique imposes across the board.

How pipeline parallelism complicates chunked prefill at scale

The basic chunking mechanism assumes, implicitly, that equal-size chunks take roughly equal time to run. That assumption holds on a single device but breaks down once a deployment spans multiple devices through pipeline parallelism. In a long sequence, later chunks attend over a progressively longer KV cache than earlier chunks do, so their attention computation costs more even though the chunk itself is the same size in tokens. Equal-size chunks end up producing unequal execution times, and in a pipeline, unequal stage times turn into bubbles, stretches where downstream devices sit idle waiting for an upstream stage to finish.

One answer to this is to resize chunks dynamically at runtime so that every chunk takes roughly the same amount of time regardless of its position in the sequence. This is workable, but it comes with its own overhead: accurate cost estimation, runtime calibration, and fine-grained fragmentation of the token stream, all of which can outweigh the bubble reduction they are meant to buy as sequences get longer.

Virtual Pipeline Parallelism, or VPP, takes a different approach: keep chunk sizes fixed, and instead restructure the pipeline layout around a V-shaped traversal of virtual stages, so that each chunk's expensive middle stages overlap with the lighter head and tail stages of its neighbors<sup>1</sup>. The growing attention cost of later chunks gets absorbed into parallel execution. VPP has been implemented in vLLM-Ascend and evaluated on three MoE-based large language models, with sequences up to 1 million tokens, across 16 Ascend 910C NPUs, and it improves throughput meaningfully over dynamic chunk resizing on both long sequences and mixed workloads.

What this reveals is a limit in the basic chunked-prefill mechanism itself: chunk size as a fixed scheduling unit is a clean idea on one device, but at scale it requires reasoning about where a chunk sits in the model's layer topology and sequence position, not just how many tokens fit in a token budget. The token-dimension abstraction that makes chunked prefill tractable has to be supplemented with pipeline-aware scheduling once the deployment spans multiple accelerators.

Chunked Prefill and MoE Models

Chunked prefill remains a sound, effective baseline for dense transformer models, where every token passes through the same fixed set of weights regardless of which chunk it belongs to. Mixture-of-experts models break that assumption. In an MoE architecture, each token activates only a sparse subset of experts, and which experts get activated can differ from one chunk to the next. Small chunks, which tight TTFT targets call for, end up loading and unloading expert weights repeatedly across iterations, and that repeated loading inflates memory traffic well beyond what a dense model would see at the same chunk size.

The MLSys 2026 oral presentation "From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill," by Gunjun Lee and colleagues, presented May 19, 2026, puts a number on this cost: chunked prefill in MoE models increases memory traffic by up to 39%, with a corresponding rise in energy consumption from redundant expert weight loads. That cost traces directly back to splitting the scheduling unit along the token dimension, which is precisely the dimension along which MoE routing varies.

The paper's proposed alternative, layered prefill, does not throw out chunking so much as change what gets chunked. Instead of splitting the token stream into windows, it splits the model's transformer layers into groups and interleaves prefill and decode vertically, across those layer groups. Because the scheduling unit is no longer tied to which tokens activate which experts, the technique sustains stall-free decoding without triggering repeated expert reloads. Layered prefill substantially reduces TTFT, end-to-end latency, and per-token energy on MoE workloads, consistently improving on chunked prefill's TTFT-versus-TBT Pareto frontier.

The lesson for practitioners is architectural. Token-dimension chunking is the right tool for dense transformers, and it remains the correct default there. Layer-dimension scheduling is the better fit once a model's routing varies by token, and production systems serving MoE models at real scale should not assume that chunked prefill's dense-model defaults carry over unchanged.

KV cache storage tier and chunked prefill's TTFT gains at long context

Every TTFT gain chunked prefill delivers rests on an assumption: that the KV cache each chunk writes can be written to, and read back from, GPU high-bandwidth memory at the speed the scheduler is built around. That assumption holds until the cache outgrows what HBM can hold. GPU HBM capacity is finite, and for long-context requests, or deployments serving many concurrent sessions, the KV cache eventually has to spill somewhere else, and everywhere else is slower by a wide margin.

The bandwidth gap between tiers is not a small discount. GPU HBM delivers on the order of terabytes per second. A CPU-to-GPU link over PCIe 5.0 delivers about 64 gigabytes per second, roughly 2% of HBM's bandwidth. Moving a large KV cache from CPU DRAM to GPU memory therefore takes an order of magnitude longer than moving the same data within HBM, and if that transfer has to complete before decode can resume, it cancels out the TTFT reduction chunked prefill was built to provide. At sufficiently long context, a large model can hold more of its memory footprint in KV cache than in its own weights, and at that point there is no avoiding a lower storage tier. The only question left is how fast that tier can serve the data.

Where the cache crosses into a lower tier depends on context length in a fairly direct way. For short contexts, a few hundred tokens, recomputing the KV state is often cheaper than storing it and retrieving it later. Once contexts stretch into the tens of thousands of tokens, storing and offloading the cache becomes the better bet on raw economics, but only if the storage tier underneath it can serve KV blocks fast enough to stay off the critical path. Chunked prefill fixes the scheduling half of TTFT. The storage tier underneath the KV cache decides whether that fix actually reaches the user, or gets absorbed by a slow transfer the scheduler never anticipated.

Sources

  1. Efficient LLM Inference via Chunked Prefills
  2. VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference
  3. Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference

More in Inference Serving