Tensor Parallelism and KV Cache Sharding Patterns
Tensor parallelism hits a hard ceiling on KV cache when it runs out of attention heads to shard.
Omar Delgado
Reporter
Omar Delgado is a reporter at Inference Engineering covering inference serving. Based in Melbourne, Omar has written for Inference Engineering since 2015.
1 story · Melbourne
Tensor parallelism hits a hard ceiling on KV cache when it runs out of attention heads to shard.