Speculative Decoding Storage and Memory Overhead
Trading latency gains for hidden memory costs that scale with context length.
Lena Petrova
Senior Writer
Lena Petrova is a senior writer at Inference Engineering covering inference serving. Based in Montréal, Lena has written for Inference Engineering since 2021.
1 story · Montréal
Trading latency gains for hidden memory costs that scale with context length.