Scheduler Design for Long-Context and Short-Context Mixed Traffic

KV cache residency, not queue position, determines how to schedule mixed-length requests fairly.

Senior Writer · · 10 min read
Cover illustration for “Scheduler Design for Long-Context and Short-Context Mixed Traffic”
Inference Serving · October 5, 2026 · 10 min read · 2,319 words

A common instinct when short requests start stalling behind long ones is to add a priority queue: push short, interactive traffic to the front, let long document jobs wait their turn. That fix treats the problem as one of ordering, and it does not work, because the actual constraint is how much KV cache capacity the GPU has free at any given moment, not which request sits where in line. Mixed-traffic scheduling is a cache residency problem before it is anything else, and every mechanism this piece covers, chunked prefill, deadline-aware priority, preemption, cluster routing, traces back to who holds cache space right now, and for how long.

Mixed-traffic scheduling as a KV cache residency problem

Continuous batching across replicated GPU instances worked well under short-context workloads because requests were roughly uniform in length and in how much cache each one needed. A scheduler could pack many of them onto a single GPU, keep the batch full, and hold utilization high without worrying much about any one request crowding out the rest. Long-context requests break that uniformity. A single long-context request on a large model can need tens of gigabytes of VRAM just to hold its own KV cache, and that footprint eats into the working memory the GPU needs to batch short requests alongside it. Effective batch size collapses, and throughput falls with it, because there is no longer room on the GPU to serve both populations at once.

The visible symptom is head-of-line blocking: short requests that arrive while a long request is in prefill end up waiting. That wait has nothing to do with queue fairness. The GPU simply has no cache space left to admit them concurrently. Research from Microsoft Research and Georgia Tech on the Medha system, published on arXiv, shows this empirically: even a small fraction of long requests mixed into a heterogeneous workload produces sharp median and tail latency spikes for the short requests sitting alongside them, under non-preemptive systems such as LoongServe. The two request types also pull on different physical resources. Short requests spend most of their time compute-bound during prefill, while long requests shift into a memory-bound regime during decode, where KV cache transfer saturates HBM bandwidth. A single undifferentiated scheduler, one built around queue position rather than cache occupancy, cannot serve both resource profiles well at the same time. Everything that follows in this piece is an attempt to design around that fact rather than against it.

PagedAttention, the memory hierarchy, and the physical limits every scheduler must respect

Specific scheduling mechanisms depend on the hardware reality underneath them, so that reality has to be precise first. KV cache lives in blocks, allocated and freed the way a block-based memory manager handles memory pages, and those blocks sit somewhere in a tiered hierarchy: GPU HBM at the top, then PCIe-connected host memory, then NVMe storage further down. Each tier crossing costs bandwidth and time, and that cost is the bandwidth cliff every scheduling decision has to respect.

Techniques that shrink the footprint of a single request help, but they do not remove the underlying problem. Grouped Query Attention, used in models like Llama 3, cuts the number of KV heads relative to query heads by a large factor, so per-token cache size drops. Quantization does similar work at a different layer: vLLM supports FP8 KV cache storage through its --kv-cache-dtype fp8 setting, and NVIDIA's Blackwell architecture adds NVFP4 support. Both reduce how much memory a given context consumes, but a long enough context still eventually crosses from HBM into slower tiers, and compression only delays that crossing rather than eliminating it.

That matters directly for scheduling. If a scheduler does not track which tier a request's KV blocks currently sit in, it will trigger reloads across PCIe or NVMe without warning, and those reloads show up as time-to-first-token spikes for every short request waiting behind them. This is no longer a software-only concern. NVIDIA announced the Inference Context Memory Storage Platform at CES 2026 and renamed it Context Memory eXtension, or CMX, at GTC in March 2026. CMX standardizes NVMe-backed KV cache offload at the cluster level, built on BlueField-4 DPUs and the NIXL transfer library. Tier-aware scheduling is becoming something the platform itself provides rather than a heuristic a serving team bolts on afterward. Each scheduling mechanism in the sections below is best read as a specific answer to the question of which blocks live where, and at what cost to move them.

Chunked prefill as the prerequisite for preemptive scheduling in mixed-traffic systems

A long prefill, left unchunked, is atomic. The GPU is committed to it from start to finish, which on an H100 can run several minutes for a very long context, and nothing else can be scheduled in that window no matter how the priority policy is written, which is the condition chunked prefill exists to break. By splitting a long prefill into smaller pieces processed across multiple scheduler iterations, chunking creates points where the GPU can be handed to another request mid-prefill. Without that capability, preemption has nothing to act on: there is no safe moment to interrupt a request that only yields at the very end of its computation. Chunking is the mechanism that makes a long request interruptible at all, and every priority or deadline policy discussed in the next section depends on that interruptibility existing first.

Two objections have historically been raised against chunking: that it causes KV cache read amplification, and that batching decode steps together with chunked prefill steps drives up decode latency. Medha's results push back on both. Chunks as small as 40 tokens reach near-optimal arithmetic efficiency under grouped-query attention, and Medha's adaptive chunking adjusts chunk size dynamically as the computational bottleneck shifts between MLP and attention operations over the course of a long prefill. Chunk size in this design is not fixed in advance. It tracks the changing character of the computation, and Stream Pipeline Parallelism pipelines those chunks across stages to hold both throughput and predictable latency steady as sizes shift.

A related insight comes from the hardware side. Work from Georgia Tech and Samsung, published on arXiv, proposes a packing-prefetch scheduling architecture that interleaves prefill and decode phases, using the HBM bandwidth left unused during compute-bound prefill to prefetch KV data needed for the decode phase. Tested on Llama3.1 models, this reduces decode latency by exploiting the same temporal slack that chunking opens up, arriving at a complementary conclusion from the memory-system side of the problem rather than the scheduler side.

Chunking is not free. Splitting a prefill across iterations means KV blocks accumulate incrementally instead of all at once, so the scheduler has to track growing block occupancy mid-flight rather than reasoning about a request's footprint as a single fixed number. That incremental accounting ties the choice of chunk size directly to the cache eviction policy sitting above it, which is the subject of the next two sections.

Deadline-aware and workload-adaptive priority policies against starvation

Chunking makes preemption possible. It does not say which request should be preempted, or when, and getting that wrong trades one failure mode for another. A naive shortest-job-first policy stops long requests from blocking short ones, but if short traffic stays heavy, long requests can be pushed back indefinitely. That is the same convoy effect in reverse: instead of short requests waiting behind a long one, long requests never get to finish.

Medha's answer is the Length-Aware Relative Slack scheduler, LARS, which computes a deadline-normalized slack value for each request based on both its remaining work and its service-level deadline. Requests get prioritized as their slack tightens, regardless of whether they are short or long, which is what keeps the policy from collapsing into pure shortest-job-first. LARS is built to be heterogeneity-aware in a specific sense: it recognizes that a request carrying 10 million tokens of context has a legitimate deadline of its own, one that has to be honored eventually even as that request repeatedly yields to shorter requests whose slack is running out faster. The throughput gains Medha reports over non-preemptive systems come from chunking and LARS working together. Neither piece on its own produces that result.

A second system, EWSJF, developed at Toga Networks (Huawei) and accepted at KDD '26 in Jeju, South Korea in August 2026, solves an adjacent part of the same problem at a different layer. Where LARS operates inside the execution loop, deciding what to run on each scheduler iteration, EWSJF sits upstream as a request-admission scheduler, running ahead of vLLM's iteration-level batching. Its grouping algorithm discovers performance-homogeneous groups of requests directly from live traffic, without needing a fixed number of clusters set in advance, which matters because production traffic mixes shift over time rather than holding still. On top of that grouping, Density-Weighted Scoring balances urgency against fairness: scores are weighted by urgency, normalized against computational cost, combined with a logarithmic fairness term across queues so no single group can monopolize admission. Bayesian Meta-Optimization then tunes both the scoring and partitioning parameters continuously from live performance feedback, in place of fixed tuning set once at deployment.

LARS and EWSJF are not competing answers to the same question. They operate at different layers, one inside execution, one at admission, and function as complementary parts of a single design rather than alternatives to choose between. What ties them together is that both are, underneath their respective mechanics, deciding who holds cache residency next and for how long, not simply who gets to run next in wall-clock terms. Even the best-designed preemption policy running on a single instance, though, only has so much cache capacity to work with. When pressure on that cache is spread across an entire cluster rather than concentrated on one machine, single-instance scheduling runs into a ceiling. Routing decisions start to matter at that point.

Preemption and the KV cache eviction cost it is meant to avoid

Preemption helps only under a specific condition: the cost of reloading or recomputing the KV cache blocks freed by evicting a long request has to be lower than the latency damage that request was causing by sitting on the GPU. At very long context lengths, if there is no tiered storage system to fall back on, that condition stops holding. When a long-context request gets preempted and its KV blocks pushed out of VRAM, those blocks have to go somewhere: either stored on a lower tier for later reload, or discarded and recomputed from scratch. Both options cost more as context length grows, so the very mechanism meant to relieve cache pressure can itself become expensive at scale.

This creates a granularity problem that chunking alone cannot solve. Preempt too often, with small chunks and frequent context switches, and reload overhead piles up. Preempt too rarely, with large chunks and infrequent yields, and the convoy effect that preemption was supposed to fix comes right back. Medha's adaptive chunking addresses this directly by sizing chunks to hold arithmetic efficiency steady while also minimizing how many eviction events are needed to make room for short requests.

A different research line sidesteps the tradeoff rather than tuning it. A different system routes long-context requests to a processing-near-memory tier once they cross a length-based threshold, and it migrates them there instead of preempting them after admission to the GPU. Length-based placement combined with runtime migration lets the system handle a request's context growing over time, and because the eviction event never happens in the first place, there is no reload or recompute penalty to pay later. This system's interconnect infrastructure, with remote-procedure-call and direct-memory-access support, lets KV blocks move between GPU and PNM tiers without the CPU sitting on the critical path as a bottleneck. On mixed-length workloads, NELSSA reports better decode throughput and lower P99 latency than GPU-only baselines, with the gains coming from larger batch sizes on GPU, made possible by offloading long-context KV state elsewhere.

The lesson for scheduler design is specific: whether a given request should be preempted depends on where its KV blocks currently sit in the storage hierarchy, and the scheduler has to know that directly rather than guess at it. Single-instance preemption policies, no matter how well tuned, run into a hard ceiling set by eviction cost, and that cost is itself set by how KV state is distributed across the hierarchy. That ceiling is also, by its nature, a cluster-wide fact, not something any one instance can solve by itself.

Load-aware and cache-affinity routing as cluster-level KV residency management

At the cluster level, the routing decision depends on where KV cache state currently lives rather than which instance has spare compute cycles. Sending a request to the wrong instance forces a cold-start KV load or a full recompute, erasing whatever careful preemption and priority work that instance's internal scheduler was doing.

Session-affinity routing is the starting point: repeat requests from the same session get sent back to the same instance, so the prefix KV blocks from earlier turns stay resident in VRAM and do not need to be recomputed. That is the straightforward cache-hit case. Cache-affinity routing extends the same idea further by dropping the assumption that cache locality only matters within a session. Instead, the router tracks which instances currently hold which KV prefix blocks and sends new requests toward whichever instance has the greatest overlap with its content, even when that request comes from a session the router has never seen before. Treated this way, KV cache residency becomes a first-class input to the router's decision, and it stands alongside compute load rather than behind it. That is the same lens this entire piece has applied at every layer: chunking decides when a long request can yield cache space, LARS and EWSJF decide who gets it next, eviction cost decides whether taking it away was worth the price, and cache-affinity routing decides, before any of that, which machine a request should land on first.

Sources

  1. No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
  2. NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
  3. EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
  4. Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories

More in Inference Serving