AI storage throughput is the rate at which a training or inference system can deliver useful data to accelerators without making GPUs wait. The useful word is important. A storage platform may advertise hundreds of gigabytes per second and still underperform when a real workload opens millions of small files, reads metadata repeatedly, writes synchronized checkpoints, or causes every worker to request the same shards simultaneously.
Within NVIDIA AI Infrastructure, storage is the first large data-movement boundary. The existing storage networking fundamentals article provides the protocol/path context; this page focuses on sizing and measuring AI workload demand.
NVIDIA GPUDirect Storage provides a supported direct DMA data path between GPU memory and local or remote storage in appropriate environments, avoiding an extra bounce through CPU memory and reducing CPU load. It improves one part of the path; it does not fix an undersized file system, metadata server, network, or dataset layout.
Start with bytes per sample and samples per second
For training, approximate per-GPU read demand from sample size × samples/second, then multiply by active GPUs and account for augmentation, compression, caching, and sharding.
For example, 8 GPUs each consuming 500 MB/s create 4 GB/s of sustained node demand before checkpoint writes or replication overhead.
Cluster design should use the expected steady and burst demand, not a generic “X GB/s per GPU” rule that ignores model/data format.
Metadata throughput can dominate small-file datasets
Image, audio, document, and scientific datasets often contain huge file counts. Opening, stat-ing, listing, and resolving paths can bottleneck before bulk data bandwidth is saturated.
Measure open/lookup latency and metadata operations per second alongside read throughput.
Sharding many small samples into larger container formats can reduce metadata pressure, but the shard size should still support parallel workers and partial reads efficiently.
Checkpoint writes create synchronized bursts
Distributed training can produce large checkpoints from many ranks at similar times, creating a write burst far above ordinary dataset-read traffic.
If every node pauses compute while writing, checkpoint duration becomes direct lost accelerator time.
Measure checkpoint wall time, aggregate write bandwidth, metadata load, and recovery-read performance; then consider staggered, asynchronous, incremental, or tiered checkpoint strategies where the framework supports them.
Read amplification makes raw dataset size misleading
Data loaders can reread records across epochs, decode/compress on CPU, shuffle indexes, or request overlapping samples.
One terabyte of source data can generate many terabytes of storage/network traffic during one job depending on caching and preprocessing.
Profile actual bytes read from storage and cache-hit ratio instead of calculating demand from dataset capacity alone.
Local NVMe and shared storage solve different phases
Shared parallel/object/file storage provides durable cluster-wide access; local NVMe can cache hot shards, stage training data, or absorb temporary checkpoint output.
Staging reduces shared-storage pressure at the cost of copy time, local capacity management, and cache invalidation.
Decide whether the workload benefits from pre-stage, read-through cache, burst buffer, or direct shared reads based on job duration and reuse.
GPUDirect Storage can remove CPU bounce buffers
Current NVIDIA GPUDirect Storage design guidance describes a DMA path between storage and GPU memory that can bypass CPU memory staging.
This can increase effective bandwidth and reduce CPU utilization and latency when the storage, network, driver, file-system, and GPU topology support the path.
Verify GDS is active rather than assuming the library name means direct I/O occurred; unsupported paths can fall back to host-staged transfers.
Storage network and GPU network may contend for the same fabric
Large clusters often use InfiniBand or Ethernet/RDMA for both distributed collectives and remote storage.
Checkpoint bursts or dataset reads can compete with NCCL traffic and increase training-step time even when storage itself is fast.
Model oversubscription, rail topology, QoS/congestion control, and peak simultaneous traffic across both data and collective flows.
NUMA and PCIe placement still matter on the storage path
Local NVMe controllers and NICs attach to specific PCIe roots and NUMA nodes. A data path can cross CPU sockets before reaching the target GPU if the process, storage/NIC, and GPU are poorly placed.
NUMA for GPU Workloads explains how topology-aware placement can remove those detours.
Measure PCIe and CPU/memory traffic when storage bandwidth plateaus below the media/network capability.
Benchmark the same access pattern the framework uses
Sequential FIO or IOR numbers are useful capacity ceilings but do not represent every AI data loader.
Test worker count, file size distribution, shard pattern, read size, concurrency, cache state, checkpoint size, and number of simultaneous jobs.
A storage platform should be benchmarked cold and warm so cache is visible rather than accidentally credited as backend throughput.
GPU idle time is the ultimate storage symptom
Track data-loader wait, GPU utilization/SM activity, storage latency, read bandwidth, CPU decode utilization, and queue depths together.
A GPU with low utilization during training could be storage-bound, CPU-preprocessing-bound, synchronization-bound, or simply running a low-parallelism kernel.
Correlate the phases: if GPU activity drops while storage reads or data-loader waits spike, storage is a likely contributor.
Storage throughput is successful when it preserves accelerator duty cycle
The mature design sizes sustained reads and burst writes, handles metadata, separates/cache-stages intelligently, validates GDS where used, accounts for shared fabric contention, and measures GPU wait directly.
The goal is not the highest benchmark number. It is keeping accelerators doing useful compute while preserving recoverable, durable data paths at cluster scale.
Dataset layout should be treated as part of storage architecture. Frameworks that read one file per sample can create huge metadata pressure, while sharded formats can improve sequentiality and reduce open/close overhead. The trade-off is that very large shards can reduce worker parallelism or force workers to read more data than needed. Test shard size against the actual number of data-loader workers, node count, and shuffle behavior.
Compression can move the bottleneck from storage to CPU. A compressed dataset reduces bytes read from disk/network, but the CPU must decode those bytes before the GPU can consume them. If decompression saturates the host, storage utilization may look low while GPUs starve. Compare compressed versus uncompressed pipeline throughput and track CPU decode time, not just storage bandwidth.
Shared cache behavior should be understood before benchmarking. Page cache, local NVMe cache, distributed cache, or object-store client caches can make repeated epochs much faster than the first epoch. Record cold-cache and warm-cache performance separately so capacity planning does not accidentally assume every job will hit a cache populated by a previous run.
Multi-job contention should be part of acceptance. One training job may achieve excellent throughput while four simultaneous jobs compete for metadata servers, network uplinks, or object-store request limits. Capacity planning should include the expected concurrency of training, checkpointing, evaluation, and inference pipelines across the cluster rather than benchmarking one hero job in isolation.
Small-object object storage can introduce request-rate and latency constraints even when aggregate backend bandwidth is high. Parallel prefetch, range requests, sharding, local caching, and data packaging can reduce per-object overhead. The important measurement is time from data-loader request to tensor ready for GPU, not the provider’s raw storage throughput figure.
Checkpoint architecture should also consider restart time. A system that writes checkpoints quickly but takes thirty minutes to read and reconstruct state after failure is not optimized for resilience. Measure both write and restore paths, include metadata/listing costs, and validate the storage tier used for emergency restart rather than only the happy-path training loop.
Network design should separate storage flow from collective flow where required. Dedicated rails, QoS classes, or sufficient nonblocking fabric can keep all-reduce traffic from colliding with dataset reads. If both use the same links, model worst-case simultaneous demand and observe congestion counters so a storage burst is not misdiagnosed as a GPU communication problem.
Storage SLOs should be expressed in workload terms such as “data-loader wait remains below 5% of step time” or “checkpoint completes within 90 seconds at 512 GPUs.” These are more useful than a generic requirement for a certain number of GB/s because they connect infrastructure performance directly to accelerator productivity.
Data-loader parallelism should be tuned rather than maximized blindly. Too few workers leave GPUs waiting; too many can overwhelm metadata servers, cause CPU context switching, exhaust file descriptors, or create competing random reads that reduce storage efficiency. Sweep worker count with step time, CPU utilization, open operations, and backend throughput visible together.
File-system striping or object-store partitioning can alter hot-spot behavior. Large training shards placed on one metadata/data target or one object prefix can create local saturation while aggregate system capacity looks healthy. Distribute data according to the storage platform’s recommended layout and verify per-target utilization during full-cluster tests.
Inference pipelines have a different storage profile from training. Model weights may load once at startup, adapters or embeddings may change more often, and cold-start latency can depend heavily on model repository throughput. Serving platforms should benchmark model load/reload and replica scale-out time separately from per-request latency.
Storage observability should retain per-client or per-job context where possible. Aggregate backend throughput can look healthy while one training cohort experiences hot metadata shards or a slow network path. Tag jobs, datasets, mount points, and storage targets so operators can move from GPU wait time to the exact file-system or object-store bottleneck instead of treating the storage cluster as one undifferentiated resource.