Prefill and Decode Disaggregation in LLM Serving
Splitting prefill and decode onto separate GPUs unlocks independent scaling for LLM serving.
Rohan Reyes
Section
1 story in Features.
Splitting prefill and decode onto separate GPUs unlocks independent scaling for LLM serving.