Speculative Decoding Storage and Memory Overhead

Trading latency gains for hidden memory costs that scale with context length.

Senior Writer · · 10 min read
Cover illustration for “Speculative Decoding Storage and Memory Overhead”
Inference Serving · September 30, 2026 · 10 min read · 2,234 words

Speculative decoding cuts per-token latency by letting a small draft model run ahead of the large target model, proposing several tokens before the target checks them all in one pass. That split between two models brings in a set of memory costs a single-model setup never has to pay. The acceptance rate, meaning how often the target model agrees with what the draft guessed, decides how much speedup you actually get: reject the draft's tokens often enough and the target ends up redoing work it would have done anyway, so the gain shrinks or disappears. The memory bill shows up whether or not the speedup does, which matters for anyone sizing hardware. The trade, then, amounts to a compute cost turned into a memory one, more than an exchange of compute for latency. It's a compute-bound cost turned into a memory-bound one, and that swap is structural. Everything below traces out what that swap actually costs, starting with the most basic layered memory costs speculative decoding leaves behind.

Draft model weights as a permanent addition to GPU memory budget

The draft model's weights have to sit in GPU memory next to the target model's weights for as long as the service runs, which turns picking a draft model into a budgeting exercise. It's tempting to think you could park the draft weights somewhere cheaper, CPU memory or an NVMe drive, and pull them in when needed. That doesn't work here. The draft model has to live on the same serving node as the target model, because the moment you push its weights out to CPU or storage, you bring back the exact latency spec decoding was built to remove. The footprint is a fixed cost that stays constant regardless of load. It's a fixed cost, sitting on the GPU around the clock, for as long as the service is live.

That leaves a spectrum, and every deployment sits somewhere on it. A very small draft model costs almost nothing in VRAM, but it tends to diverge from the target model's output distribution more, meaning lower acceptance rates and less of the promised speedup. A larger draft model tracks the target more closely and gets accepted more often, but it eats a bigger, permanent slice of the memory budget to do it. There's no universal right answer here, only a trade a team has to make deliberately, against its own hardware.

One more constraint narrows the field before quality even enters the conversation: the draft model has to share the target model's tokenizer and vocabulary. Mismatched vocabularies force token remapping, and that remapping is itself overhead nobody wants sitting in the hot path. In practice, this keeps most draft model choices inside the same model family as the target, which limits how far apart the two can be sized in the first place.

KV cache duplication across speculative paths

If the draft model's weights are a fixed cost, the KV cache is where the real volatility lives. Standard speculative decoding keeps separate cache entries for the draft model, the target model, and every speculative path under review at once, though some designs share or drop the drafter-side cache to ease the load, and this is what actually dominates runtime memory overhead once the system is serving traffic. Single-model serving sizes its KV cache for one model across the active batch, full stop. Speculative decoding breaks that simple accounting into pieces. The draft model builds its own cache during the speculation phase, one entry per candidate sequence, and none of it can be released until verification says whether the guess held up. The target model builds a separate cache during verification: accepted tokens get folded into the committed cache, while rejected tokens force a discard and a rollback of whatever the draft had staged.

Picture the moment verification actually runs. At that instant, the target model's cache from the current step, the draft model's cache for its whole proposed run, and the committed cache from everything accepted so far are all live in GPU memory simultaneously. That overlap, not the average load, is the peak, and it's the number infrastructure teams need to design around, not the steady state.

Generate more draft tokens per cycle, and you raise the odds of stringing together a long run of accepted tokens. That run of accepted tokens is where the throughput gains actually come from when acceptance is running high. But every one of those extra draft tokens is sitting in memory as speculative cache, and if the run gets rejected, that memory was spent for nothing. When acceptance rates are low, a long window isn't an investment, it's a standing loss. Window length and acceptance rate aren't variables you can tune in isolation, they move together, and treating them separately is how teams end up over-provisioned or starved.

Tree-based decoding pushes this further still. Instead of one linear guess, the draft model branches into several candidate paths, and the target model has to verify all of them in the same pass. Every branch carries its own cache subtree that has to be tracked and evaluated at once. Offering the target model more options this way can lift the acceptance rate, but the cache footprint doesn't grow in step with it, it grows combinatorially with the number of branches and how deep each one runs. Whatever memory manager sits underneath has to know which subtrees got accepted and clear out the rejected ones fast, or cache fragmentation starts eating into headroom nobody budgeted for. Speculation window length, defined as the number of draft tokens generated before verification, multiplies KV cache pressure.

Verification buffers and the bandwidth cost of the reject-rewind cycle

There's a cost here that has nothing to do with how much memory is allocated and everything to do with how fast that memory has to move. Every rejected token triggers a cache rewind, and that rewind eats memory bandwidth and holds up the next draft cycle until it's done, which is how a swing in acceptance rate turns into a latency spike a user actually feels. Accepted tokens get appended onto the committed cache, and rejected ones force a rollback to the last point that was actually confirmed. That rollback isn't a bookkeeping formality. It means re-reading and re-writing cache state in HBM, bandwidth that would otherwise be feeding the next forward pass through the model.

This is the crux of the whole section: during verification, HBM bandwidth, not compute, is what's actually scarce. The forward pass at this step isn't compute-bound the way a large training batch would be, because the batch here is small and the cache reads are what dominate the time it takes. Whatever bandwidth is provisioned for cache access sets a hard ceiling on how fast the system can recover from a rejected token and get the next draft cycle moving.

That ceiling gets tested hardest when rejections cluster together. A long run of bad guesses in a row means a rewind for every single one of them, stacked back to back, each one blocking the next draft cycle from starting. Acceptance rate variance is common enough on its own, especially once the input drifts away from whatever distribution the draft model was trained on, and when that happens, the bandwidth cost becomes a systemic bottleneck rather than an occasional tax.

How context length amplifies every overhead category

Long-context serving is exactly where speculative decoding's latency gains look most attractive on paper, and it's also exactly where the three overhead categories already laid out, weights, KV cache, and verification bandwidth, stack up worst. None of these costs stay flat as context grows. They multiply against each other.

Start with the cache. Its size scales linearly with sequence length for both models, so doubling the context window doubles the cache footprint for the draft model and doubles it again, independently, for the target. Pushing context out to the tens or hundreds of thousands of tokens makes the target model's cache alone outsize the model's own weights in GPU memory. Stacking the draft model's cache on top of that makes total VRAM pressure severe fast. Tree-based speculation is especially exposed here: each branch now carries a long cache of its own, and holding several of those branches at once can push the whole approach past what most operators can actually fit in memory.

Bandwidth scales right along with it. Longer sequences mean bigger cache reads on every verification pass, and every rollback after a rejection now has to touch more data than it would at a shorter context. There's a real ceiling here: past a certain context length, the bandwidth spent on verification and rollback starts eating into the very latency gains that made speculative decoding worth adopting. The picture isn't uniform. Research on long-sequence speculative decoding at large batch sizes has shown real speedups are still achievable, so the ceiling depends heavily on batch size and workload shape, not context length alone. Still, the direction of the pressure is clear enough that teams pairing speculative decoding with long-context products in their roadmap need to size memory and bandwidth for how the two interact together, not for each one on its own.

Sizing GPU memory and storage bandwidth for a speculative decoding deployment

Sizing a speculative decoding deployment calls for a different methodology than sizing single-model serving. The number that matters is peak co-resident memory, meaning both models' weights plus both KV caches at the exact moment verification runs, not whatever the average utilization graph shows over an hour.

Start with weights. Adding up the target model's footprint and the draft model's footprint at whatever precision the deployment actually serves at, BF16 or FP8 or otherwise, produces a sum that should be treated as a floor that doesn't move, not something that shrinks when load drops. If FP8 quantization is being used on the draft model to trim that footprint, remember the dequantization step during verification isn't free either, it costs its own slice of memory bandwidth.

Then size the cache for its peak, not its average: the moment both draft and target caches are fully built out, at the longest speculation window and the longest sequence length the batch will see. That number runs well past what a single-model deployment would need at the same sequence length and batch size. For tree decoding, multiply again by the branch count, which in practice tends to confine tree decoding to shorter contexts or to setups running aggressive cache compression.

Compression is a real lever here, not a nice-to-have. Compressed cache can be pushed out to NVMe and streamed back for context that's already been committed, but the live speculative window is a different story: its latency requirements are tight enough that offloading it isn't realistic. Drawing that line, what stays in GPU HBM versus what can move to NVMe or host memory, is the central storage design decision in any speculative decoding system.

Storage bandwidth has to keep pace with that line too. The rate at which accepted cache streams out to NVMe, or streams back in for prefix reuse, has to match serving throughput, or a queue builds up and starts stalling the GPU. High-throughput NVMe with GPUDirect Storage keeps the CPU out of that path entirely, which cuts the latency penalty tiering would otherwise add. Where cache needs to move across nodes, for prefix caching or disaggregated serving setups, an RDMA-capable storage fabric becomes part of the same decision.

Speculation window length is the last dial, and it should stay a dial that operators set deliberately rather than a default pulled from a config file. Shorter windows bring peak cache pressure down at the cost of some of the potential speedup, and the right setting depends on measured acceptance rate and whatever VRAM headroom actually exists.

Operational signals that indicate speculative decoding overhead is the binding constraint

Every overhead category traced through this piece leaves a specific fingerprint in production, and knowing which one is flashing tells an operator whether the fix is more memory, a shorter window, or more storage bandwidth, not more GPUs for raw compute.

On the memory side, watch for effective batch size shrinking under speculative decoding compared to single-model serving at the same sequence length: that's the system trading away concurrency to make room for speculative throughput. That's the co-resident weights and dual cache footprint showing its cost directly.

On the bandwidth side, latency variance that spikes specifically during stretches of high rejection rates is the rewind cycle eating HBM bandwidth that should be going to forward passes. HBM utilization staying high while compute utilization sits moderate is the same story told a different way: memory bandwidth, not FLOPS, is what's actually constrained.

On the storage side, an offload queue depth that keeps growing under sustained load means storage bandwidth isn't keeping up with how fast accepted cache needs to be written out. Latency that degrades specifically on long-context requests, where cache has to be reloaded from NVMe, points at storage read bandwidth as the bottleneck.

Acceptance rate is the earliest warning available, producing signals that predict trouble before other metrics move. Tracking it per request, a falling trend predicts which of these symptoms is about to get worse before any of the system-level metrics move. It's the cheapest instrumentation in the whole stack, and it's the one number that tells you the others are about to turn.

Sources

  1. H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache
  2. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
  3. Resource-Efficient Speculative Decoding for Long-Context LLM Serving
  4. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
  5. Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache
  6. Memory overhead | Speculative Decoding Intermediate Course | The Neural Base

More in Inference Serving