GPUDirect Storage Integration in Inference Pipelines

GDS excels at training but underperforms for inference's small random reads.

Senior Writer · · 9 min read
Cover illustration for “GPUDirect Storage Integration in Inference Pipelines”
Inference Serving · October 3, 2026 · 9 min read · 2,106 words

GPUDirect Storage was purpose-built to move large, sequential blocks of data across GPU workloads, and training is where that design assumption holds most cleanly. Understanding what GDS was actually solving matters before anyone wires it into an inference pipeline, because GDS was purpose-built for large-block sequential I/O across GPU workloads, including training. The traditional path for a checkpoint write traverses five hops: GPU HBM, PCIe, CPU pinned memory, the OS kernel buffer, and the NVMe driver queue. GDS collapses the middle three of those hops into a direct route, GPU HBM to PCIe bus to NVMe controller, using the cuFile API and the nvidia-fs kernel driver.

For checkpointing at training scale, that collapsed path produces a clear speedup. A 100-billion-parameter model produces checkpoints that range from hundreds of gigabytes to over a terabyte, and large-scale training clusters need to write more than a hundred checkpoints a day. If those clusters are going to complete each one inside their overhead budget, that's only realistic with the direct path GDS opens. That is the problem GDS was engineered to solve: sustained, large-block, sequential movement of data where throughput, not responsiveness, is the governing constraint.

Inference runs on the opposite I/O profile. Where training reads and writes in large sequential sweeps, inference generates small, random, read-heavy access patterns, and it does so under SLA constraints that prize responsiveness over raw throughput. A study tested a large number of I/O configurations across pre-training, fine-tuning, and inference phases and found that io_uring achieves the lowest latency and competitive IOPS for small random I/O on NVMe during inference, while GDS excels in the coarse-grained sequential reads and writes that dominate training. The lesson from that research is that the choice of datapath needs to match the phase of the pipeline, and applying GDS unmodified to small-block random reads leaves real performance on the table.

Model Loading and the Cold-Start Problem

Model loading is the one inference scenario that looks structurally like training I/O: large, sequential reads of weight data, rather than the small random access that defines serving traffic. That similarity is why GDS delivers its full payoff here, even though it underperforms elsewhere in the inference pipeline.

When you load a large model the conventional, CPU-mediated path, it runs through several serial steps. The checkpoint is read from storage into CPU system memory, the weights are deserialized, they are optionally quantized on the CPU, and the result is copied to each GPU over PCIe one at a time. That sequence is largely single-threaded, sequential, and CPU-bound. The two are not competing solutions to the same problem; they operate at different layers of the loading pipeline, and the gain from one does not substitute for the gain from the other.

Cold-start latency is not a minor inconvenience to be smoothed over. It directly limits how fast autoscaling can respond to demand, how quickly a cluster recovers from a fault, and how efficiently compute resources get spent, because every unit of compute time spent loading a model is a unit not spent serving requests. The architectural fix that GDS enables is pre-sharding the checkpoint across tensor-parallel ranks at the filesystem level, so that all eight GPUs read their own shards in parallel, directly into HBM, over EFA, with CPU memory bypassed entirely. On AWS infrastructure, FSx for Lustre combined with GDS uses eight or more of the P5en instance's 16 EFA interfaces for this direct storage-to-GPU transfer, drawing on a portion of the interface count rather than the instance's full aggregate network bandwidth.

FP8 pre-quantization compounds the gain further. FP8 pre-quantization is a force multiplier here: it halves the checkpoint size of a large frontier model, halving load time before GDS even enters the picture at all, and combined with GDS's parallel loading path, the effect compounds.

KV Cache as the Dominant Storage Problem in Production Inference

Model loading is a bounded, one-time event at the start of a deployment's life. KV cache management is continuous, running for the life of every session a production system serves, and at the context lengths modern deployments operate at, it is not an optimization nice-to-have but a physical constraint that makes NVMe tiering necessary.

Production serving stacks now run a four-tier memory hierarchy to manage this constraint. VRAM holds the hot blocks of active sessions, CPU DRAM extends the cache at microsecond access times, local NVMe holds cold sessions and long-tail prefixes, and where GDS-class paths are in place, reloading from NVMe bypasses the CPU entirely.

Offloading KV cache to this hierarchy earns its keep in specific circumstances. Multi-team retrieval-augmented generation workloads, or multi-codebase agentic workloads, cycle many distinct prefix sets through eviction and benefit from a deeper cache tier to hold them. Long-context, multi-turn conversations at extended token lengths reach a point where the cost of re-running prefill exceeds the cost of transferring stored KV over RDMA, favoring offload. Cross-node session migration benefits as well, since the receiving GPU can resume decode directly from stored KV instead of re-running the entire prefill computation.

But the same technique backfires in a different setting, and you need to take that boundary seriously rather than gloss over it. General-purpose chatbot deployments, where each request is short and context reuse across sessions is low, do not benefit from KV offloading: the I/O overhead of writing to and reading from a lower tier exceeds whatever prefill savings offload might otherwise provide. The claim for KV offload holds for workloads where prefix sharing and session length are long enough to justify the round-trip cost, and it does not hold as a general prescription for every inference deployment.

A Springer Journal of Supercomputing paper from 2025 gives a sense of the ceiling available when the access pattern suits a direct SSD-to-GPU path. GPU-centric information retrieval backed by GDS, using dynamic embedding computation in a system called ESPN-LIVE, cuts query latency by a substantial multiple and reduces storage costs by up to 16 times. That result illustrates what is achievable when the access pattern is retrieval-shaped, which is a narrower condition than general inference serving but a meaningful one for the workloads that fit it.

NVIDIA's ICMSP and SCADA: what the platform stack now provides natively

NVIDIA's ICMSP (also called CMX) and SCADA represent two distinct layers of the solution: ICMSP extends the context memory address space into NVMe to address KV capacity, while SCADA addresses the control-path gap that GDS left unsolved.

ICMSP, announced at CES 2026, extends GPU KV cache into NVMe-based storage. It relies on the Rubin GPU cluster for cache capacity and on the BlueField-4 DPU for hardware-accelerated cache placement, which eliminates metadata overhead that would otherwise slow the path down. NVIDIA's Dynamo KV cache offload engine, paired with NIXL (the NVIDIA Inference Xfer Library), provides accelerated KV cache sharing across AI nodes. NVIDIA states the platform delivers up to 5 times greater power efficiency than traditional storage and up to 5 times higher tokens-per-second throughput.

Data movement within ICMSP follows the same semantics GDS established. Local NVMe transfers use PCIe peer-to-peer communication, and remote storage transfers use NVMe-over-Fabrics, NFS-over-RDMA, or RDMA-capable distributed file systems, with KV data landing directly in GPU memory in every case, never transiting host RAM.

SCADA addresses a different part of the problem: eliminating the CPU's remaining role in the control path. Introl's analysis of SC25, held in November 2025, reports that Micron demonstrated SCADA on an H3 Platform Falcon 6048 server fitted with 44 Micron 9650 PCIe Gen6 SSDs, recording very high IOPS figures for 512-byte random reads, with each SSD delivering millions of IOPS and the array scaling linearly as SSDs were added. SCADA achieved a significant speedup over CPU-initiated storage access in that demonstration, and it reduced hardware costs substantially compared to holding the equivalent data entirely in DRAM.

GPU memory hierarchy constraints that determine when NVMe tiering is feasible

The decision to tier KV cache out to NVMe is a direct function of how much HBM capacity a GPU has relative to the model's size and the context length being served.

LLM decode is memory-bandwidth-bound rather than compute-bound, which makes HBM bandwidth one of the most consequential specifications for inference performance, since the decode phase lives or dies on how fast memory can be read. Current-generation hardware sets the ceiling here. The H100 SXM5 delivers 3.35 TB/s of HBM bandwidth. The H200 SXM5 delivers 4.8 TB/s with 141 GB of HBM3e. The B200 SXM6 delivers 8 TB/s with 192 GB of HBM3e, and an 8-GPU DGX B200 system provides 1.5 TB of total GPU memory across the cluster.

Those numbers matter directly against the workload cases described above. Even with the HBM headroom current-generation GPUs offer, very long context windows at BF16 precision exceed what a single GPU can hold for a single user. NVMe tiering therefore remains necessary for frontier context lengths regardless of GPU generation. The crossover point sits where quantization and HBM capacity intersect: for context lengths where aggressive quantization keeps the KV cache inside available HBM, adding NVMe tiering only adds latency with no offsetting benefit. For context lengths or concurrency levels that exhaust even compressed KV headroom, GDS-class NVMe paths stop being optional and become the only way to serve the request at all. When available HBM is exhausted, the GPU memory hierarchy becomes the binding constraint on whether a given request can be served.

Network fabric requirements for GDS to work at inference scale

Everything described so far, from model loading to ICMSP's KV tiering, depends on a network fabric capable of carrying it, and that fabric is a prerequisite rather than an enhancement. GDS running over NVMe-oF requires RDMA and does not function over standard TCP/Ethernet, which makes the fabric a hard dependency for distributed KV cache tiering rather than a layer that can be added later.

Running NVMe-oF with GDS calls for an InfiniBand or RoCE network built on GPUDirect RDMA-capable NICs, and full GDS performance depends on that fabric being in place. NVMe-oF over TCP is supported, but it does not deliver the same latency profile, and multi-node tensor parallelism needs high-bandwidth InfiniBand or RoCE with GPUDirect RDMA to perform as intended. That dependency is verifiable directly in NCCL logs, which show NET/IB/GDRDMA as the active transport where the fabric is correctly configured, rather than falling back to NET/Socket. NVMe-oF can still improve meaningfully on legacy NFS or iSCSI deployments, but neither of those older protocols delivers the latency profile that KV offload needs to meet inference SLA requirements.

Production deployments already reflect this fabric-first ordering. Introl's analysis notes that NVMe-oF adoption is growing quickly as organizations extend PCIe-level latency characteristics across the network to keep GPU utilization high in distributed AI clusters. AWS's P5en configuration, paired with FSx for Lustre, offers a comparable picture: 16 EFA interfaces at 200 Gbps each, for 3,200 Gbps of theoretical aggregate bandwidth, with FSx for Lustre drawing on eight or more of those interfaces specifically for direct storage-to-GPU transfer.

For an on-premises deployment, this ordering has a practical consequence: the network fabric decision has to precede the storage decision.

Lossless compression for FP8 and BF16 KV blocks in production

Compressing KV blocks before they move to a lower tier, whether that tier is CPU DRAM or NVMe, reduces the volume of data that has to cross PCIe or the RDMA fabric for a given reload. The gain compounds with the same logic that applies to pre-sharded checkpoint loading over EFA: a smaller payload takes less time to move across any fixed-bandwidth path.

The practical weight of this layer depends on where it sits in the pipeline. If compression runs before data leaves HBM, it adds compute overhead to the decode-bound GPU, already the tightest resource in the serving pipeline, so it only pays off when the bandwidth you save on the transfer exceeds the cycles you spend compressing. Compression applied at the storage tier itself, after data has already left the GPU, avoids that trade-off but depends on the storage layer's own compute capacity to do the work, which is the kind of hardware-accelerated placement that BlueField-4 style DPUs are built to absorb without adding latency back into the critical path. Either way, the same principle that governs the rest of this stack applies here too: the benefit of compression is a function of context length, concurrency, and how often a given KV block gets reloaded. Short sessions with little reuse gain almost nothing from compressing blocks that will be evicted before they are read again. Long-context, high-reuse workloads, the same workloads that justify KV offload to NVMe in the first place, are exactly where shrinking the payload pays for itself on every reload.

Sources

  1. Towards Scalable Storage Architectures for GPU Clusters ...
  2. AI-Optimized Storage: NVMe-oF, GPUDirect & Parallel ... - Introl
  3. Accelerate LLM model loading and increase context windows with GPUDirect on Amazon FSx for Lustre and TurboQuant
  4. Storage access optimization for efficient GPU-centric information retrieval
  5. How to save GPU memory in LLM serving: Principles and operating conditions of KV cache offloading

More in Inference Serving