Tensor Parallelism and KV Cache Sharding Patterns
Tensor parallelism hits a hard ceiling on KV cache when it runs out of attention heads to shard.

Tensor parallelism restructures how KV cache gets allocated, replicated, and moved across devices. It does not just decide how weights get divided. Tensor parallelism restructures how KV cache is allocated, replicated, and moved across devices, not merely how weights are divided. Unlike a weight matrix, the cache is stateful and tied to a specific request, and it grows with context length and batch size rather than sitting fixed at some size determined at load time. Because of that difference, sharding weights across more GPUs does not automatically shard the cache in proportion, and past a certain TP width the cache stops distributing and starts duplicating instead.
That distinction sets up the real question behind this piece: given a fixed model and a fixed hardware budget, what does TP degree actually do to the KV cache, and where does that behavior stop matching intuition?
The structural ceiling TP hits inside the attention layer
Attention heads are the unit TP actually shards. Each device takes responsibility for a subset of heads, and the cache belonging to those heads has to live on the device that owns them. Every additional GPU past that line doesn't get a smaller slice of cache: it gets a complete copy. Once TP degree exceeds the KV-head count, there are no more heads left to shard, so the cache can't be divided any further.
This ceiling arrives earlier than it used to. Grouped Query Attention, now standard in a lot of large models, has pushed KV-head counts down to eight or fewer, so the point where TP stops shrinking cache per device occurs at a low TP degree even in sizable deployments. GPU count also has to divide the attention head count evenly, so a team can't just pick an arbitrary TP width and expect the model to cooperate.
Multi-Head Latent Attention, the mechanism behind DeepSeek-V2 and V3, runs into a related but distinct wall. Under standard TP, every device has to load the full latent vector c_KV regardless of how wide the TP degree gets, so scaling TP buys zero per-device cache relief. That's a different failure mode than the GQA case, not a milder version of it, and it matters because the fix for one doesn't transfer cleanly to the other, and the TP degree at which cache stops shrinking shows up sooner than most sizing spreadsheets assume, with exactly where depending on which attention variant a given model uses.
Context length turns a manageable ceiling into an existential constraint
None of this matters much at short context lengths. A KV-head ceiling that forces some cache duplication across a few extra GPUs is an inefficiency, not a crisis, when contexts run a few thousand tokens. Cache size scales linearly with both context length and batch size, so that same ceiling turns into something closer to an existential constraint once context windows stretch into six-figure token counts. At that scale, storing the KV cache for even a handful of concurrent long-context users on one GPU becomes physically impossible, and the cache alone, before a single weight tensor gets loaded, can consume the entire device's memory.
Capacity isn't the only casualty. Every decode step means reading the accumulated KV cache back out during self-attention, and that read gets more expensive as the context grows, so latency degrades even in cases where the memory technically fits. Dropping the batch size is the obvious lever to pull when cache won't fit, but it doesn't touch the read-time cost of a long history. It just trades throughput away for latency without actually solving either problem.
Teams planning hardware tend to miss the consequence for how they configure it. A TP configuration that works well for a short-context chatbot workload can be structurally wrong for a long-context workload running on the identical hardware. That decision has to get made before deployment, based on what the context length actually will be in production, not adjusted after the fact once latency numbers come back ugly, because once the ceiling from the previous section combines with six-figure context windows, whether a deployment works depends on that variable, not on a footnote in a capacity plan.
Adding GPUs versus compressing the cache
Once a team hits this wall, two levers present themselves: add more GPUs, or compress the cache that's already there. They look like substitutes for each other, and they get treated that way in a lot of capacity planning, but they solve different problems. Across Llama-2 at two model scales, three GPU types, and every matched level of memory relief, compression is cheaper by a factor ranging from roughly one-and-a-quarter to two times, and the gap widens as the required relief deepens. Compression is the only one that multiplies how much cache capacity you get per dollar spent. Treating them as interchangeable is how hardware budgets get misallocated.
The cost math backs this up in a specific way. Across Llama-2 at two model sizes, tested on three different GPU types, at every matched level of memory relief, compression comes out cheaper by a factor of roughly one-and-a-quarter to two, and that gap gets wider the more relief you need. So on a pure cost basis, compression usually wins. But compression carries a latency cost that TP doesn't: it worsens per-token latency because of contention introduced during batching, while TP moves latency in the opposite direction, reducing it. The two levers pull against each other on the latency axis, so the right choice comes down to which constraint actually binds: cost-per-token or time-per-token.
Model size draws a hard line through this decision. Below roughly 36B parameters on an 80GB device, compression dominates and extra GPUs are largely wasted spend. Above that line, TP stops being optional.
Data parallelism doesn't offer an escape hatch here, either. Every replica under data parallelism needs its own full copy of the weights and its own separate KV cache, so the main cost is memory duplication, not memory relief. That means data parallelism does nothing for a model that already exceeds device memory on its own. The natural pushback at this point is to ask why not just quantize everything and stay on a single GPU. Compression only touches KV tensors. It leaves weights untouched, so above the weight-memory wall, compression can't substitute for TP. It can only work alongside it.
KV cache quantization formats and the memory hierarchy
The format a KV tensor gets stored in decides more than its footprint on disk or in memory. It decides which tier of the memory hierarchy can hold it at all, and how much bandwidth that tier needs to deliver the cache at decode speed. BF16 remains the uncompressed reference point every other format gets measured against. FP8 is a production default in systems like vLLM and LMCache, and NVFP4 stores tensors 75% smaller than BF16, with Blackwell hardware providing native support.
Lossy quantization comes with a real cost: it trims memory footprint at the price of accuracy risk, and that trade stops being acceptable the moment KV tensors need to move between nodes without degrading. That's the specific case the SplitZip paper (arXiv:2605.01708) targets, using lossless fixed-length exponent coding on native formats like BF16 and FP8, built for transferring KV cache between prefill and decode nodes where quantization error can't be tolerated. DFloat11, presented at NeurIPS 2025, takes a related but separate approach, using dynamic-length float encoding to claim lossless compression at a reduced size, and it represents one of the more active research threads in lossless alternatives to quantization right now.
The memory hierarchy itself is expanding downward. FP8 and PagedAttention cover the in-GPU optimizations, but NVMe extends cache capacity past what CPU DRAM alone can hold. DUAL-BLADE (arXiv:2604.26557) demonstrates direct GPU-to-NVMe paths via io_uring_cmd, bypassing kernel I/O stacks for lower-latency eviction. A separate paper (arXiv:2609.11744) characterizes what external KV caching on NVMe SSDs looks like for vLLM specifically, and OSDI 2026 work titled "No buffer, no bottleneck" removes the intermediate GPU HBM staging buffer from the KV offloading path to CPU memory. CXL is emerging as another tier sitting between HBM and NVMe: Beluga, presented at SIGMOD 2026, proposes a CXL-based architecture for managing KV cache, and CXL-SpecKV (FPGA '26) uses CXL's high-bandwidth interconnects paired with FPGA acceleration to prefetch cache speculatively, cutting GPU memory pressure and reducing latency stalls.
None of these choices are interchangeable, and none of them are commodity decisions made after the fact. Format, quantization scheme, and tier placement together decide whether a given TP configuration can hit its target latency at all. A KV budget that fits comfortably on one tier at FP8 may not fit at BF16, and moving between tiers changes how much bandwidth is available to serve the cache in the first place.
Prefill–decode disaggregation and the KV transfer problem it creates
Prefill and decode want different things from hardware. Prefill is compute-bound, decode is bound by memory bandwidth, and running both phases on the same GPU means tuning for one profile while the other one just has to absorb whatever's left. Under real load, prefill interferes with decode and neither phase gets the configuration it actually needs. Splitting the two phases onto separate hardware resolves that mismatch, and both DistServe and Splitwise show meaningfully higher request throughput at the same service-level targets once prefill and decode run apart.
That fix doesn't come free. Disaggregation solves the compute-bandwidth mismatch but immediately turns KV cache transfer into a distributed systems problem that someone now has to design and own explicitly. Before disaggregation, the cache was transient state living inside a single request on a single machine. After it, the cache becomes a handoff between two separate systems, and managing it, transferring it, routing it, storing it, and expiring it across machines, becomes a core piece of the architecture rather than an afterthought.
Interconnect speed is what actually constrains how far this can scale. TP needs very fast interconnects and stays confined to single-node boundaries because of it, so scaling across nodes requires pipeline parallelism on top, and any KV transfer between disaggregated prefill and decode nodes has to travel over whatever interconnect happens to connect them. Disaggregation isn't a universal win, either. For workloads with short, concentrated outputs at small batch sizes, the overhead of transferring the cache outweighs the benefit of splitting the phases, and the right configuration depends on model size, traffic patterns, latency targets, and the hardware actually available.
Mooncake, built by Moonshot AI, tackles the routing side of this problem directly with a KV-cache-centric architecture that offloads hierarchically across GPU HBM, CPU DRAM, and SSD, routing KV blocks from prefill workers to decode workers through a distributed store instead of broadcasting from a single source. This is one concrete answer to the transfer problem disaggregation creates, and it reflects how much infrastructure now sits between prefill and decode that didn't need to exist when both phases ran on the same chip.
Sharding patterns that break through the KV-head ceiling
Three structural approaches break past the KV-head ceiling described earlier, and they get there by changing what the cache actually shards along, not by finding a clever way around the head-count limit.
Sequence-dimension sharding, also called context parallelism, partitions the input sequence itself across devices instead of partitioning attention heads. Each GPU handles a slice of the token sequence, so cache size on each device scales with sequence length rather than head count, which removes the head-count ceiling completely. The cost appears instead in cross-device communication needed to compute attention across those sequence slices. Vertumnus builds on this by treating context-parallelism degree as something the scheduler actively decides on, rather than a fixed setting, routing requests to workers running different CP degrees based on a placement cost that weighs predicted queuing delay, cache-aware prefill time, and GPU-time cost together. On a 64-GPU cluster, that approach cut mean time-to-first-token by up to 28.1% and improved token-weighted SLO attainment by a meaningful margin over the strongest baseline tested.
NVIDIA's Helix Parallelism takes a different route: decoupling the sharding strategy for attention from the sharding strategy for the feed-forward network within each transformer layer. Implemented in TensorRT-LLM for GB300 NVL72, Helix shards the KV cache, which can run to multiple millions of tokens, along the sequence dimension, while applying tensor parallelism specifically across attention heads, which keeps the cache from getting duplicated across GPUs. Tested against DeepSeek-R1 on GB300 NVL72, Helix delivers a substantial jump in both concurrency and interactivity at very long sequence lengths, and it's built specifically around what Blackwell hardware offers: the high-bandwidth NVLink domain and FP4 compute.
The third approach addresses MLA specifically. Tensor-Parallel Latent Attention, or TPLA, partitions both the latent representation and each head's input dimension across devices, runs attention independently on each shard, then recombines results with an all-reduce operation. That design directly answers the MLA failure mode described earlier, where standard TP forced every device to load the full latent vector c_KV regardless of TP degree. TPLA keeps the compressed-cache benefit MLA is built for, while still letting TP actually reduce per-device memory the way it's supposed to.
None of these three approaches erases the underlying trade-offs covered earlier in this piece: cost against latency, compression against weight-memory limits, single-node interconnect speed against cross-node scaling. What they do is give practitioners a real choice about which axis to shard along, rather than accepting head count as a hard ceiling on how far tensor parallelism can take a deployment.
Sources
- Adaptive Context Parallelism for Production LLM Serving
- More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
- D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
- Data, tensor, pipeline, expert and hybrid parallelisms | LLM Inference Handbook
- TensorRT-LLM/docs/source/blogs/tech_blog/blog22_Helix_Parallelism_Scaling_Multi_Million_Token_Decoding_with_KV_Cache_Sharding.md at main · NVIDIA/TensorRT-LLM
- KV Cache Optimization: Serve 10x More Users per GPU (2026) | Spheron Blog


