Scheduler Design for Long-Context and Short-Context Mixed Traffic
KV cache residency, not queue position, determines how to schedule mixed-length requests fairly.
Tobias Wrenfeld
Staff Writer
Tobias Wrenfeld came up through HPC networking before pivoting to AI infrastructure journalism in 2018, where he developed a reputation for dissecting interconnect bottlenecks that PR decks quietly omit. He holds an M.S. in electrical engineering from TU Dresden and has tested fabric topologies at several hyperscaler labs.
4 stories
KV cache residency, not queue position, determines how to schedule mixed-length requests fairly.
GDS excels at training but underperforms for inference's small random reads.
Tensor parallelism hits a hard ceiling on KV cache when it runs out of attention heads to shard.
Trading latency gains for hidden memory costs that scale with context length.