NUMA matters for GPU workloads because CPUs, system memory, GPUs, NICs, and NVMe controllers are physically attached to specific sockets and PCIe root complexes. A process can run correctly while repeatedly crossing the inter-socket link to reach memory or a device, adding latency and consuming bandwidth before data even reaches the GPU.
Within NVIDIA AI Infrastructure, NUMA is the locality layer beneath GPUDirect, storage, scheduling, and multi-GPU performance. NVIDIA nvidia-smi topo -m reports GPU/GPU and GPU/NIC topology plus CPU and memory affinity, making it a useful first map of the server.
Topology-aware placement is increasingly important because modern GPU, NIC, and HBM bandwidth can exceed what an accidental cross-socket path can sustain efficiently.
Start with the physical topology map
Use nvidia-smi topo -m, lspci, NUMA tools such as numactl --hardware, and system inventory to map each GPU and NIC to its closest CPU/memory node.
The NVIDIA topology matrix distinguishes local PCIe relationships from paths that traverse a NUMA interconnect such as UPI.
Store the expected topology per server SKU so swapped risers or BIOS changes can be detected during burn-in.
CPU threads should run near the GPU they feed
Data loader, framework launcher, communication progress, preprocessing, and inference server threads can all touch GPU-bound data.
Pinning CPU work to cores on the GPU’s local NUMA node reduces cross-socket memory access and interrupt movement.
Do not pin every process blindly; leave capacity for kernel, network, storage, and management work and validate actual CPU utilization.
Pinned host memory should be allocated on the right NUMA node
CUDA pinned memory improves host-device transfer bandwidth, but allocation locality still matters in multi-socket systems.
If a CPU thread allocates pinned buffers from remote NUMA memory, the PCIe DMA path may pull data across the socket interconnect before reaching the GPU.
Bind memory allocation or initialize buffers from the intended NUMA node, then measure PCIe throughput to confirm locality improved.
NIC affinity is critical for GPUDirect RDMA
A GPU and RDMA NIC that share a favorable PCIe root/topology can exchange data more efficiently than devices separated by the system interconnect.
NVIDIA GPUDirect RDMA depends on compatible peer-device topology and software support.
Use the NVIDIA topology output to select local NICs for ranks/GPUs rather than round-robin assigning interfaces by name.
Storage controllers can have the same locality problem
Local NVMe drives and storage NICs attach to one socket as well.
A training pipeline can become CPU/NUMA-bound if data comes from an NVMe drive under socket 0 but the GPU and data-loader CPU live under socket 1.
AI Storage Throughput should be measured with the final GPU/CPU/storage affinity, not with a standalone disk benchmark.
Multi-GPU jobs should group ranks by topology
Distributed frameworks assign local ranks to GPUs; the mapping should reflect NVLink/NVSwitch and PCIe/NUMA relationships where possible.
Communication libraries such as NCCL perform topology-aware planning, but process CPU affinity and NIC selection still influence end-to-end behavior.
Use per-rank environment/affinity and validate with collective benchmarks across the exact rank layout production will use.
Container orchestration can accidentally erase topology awareness
Kubernetes requests a GPU resource, but generic CPU/memory scheduling does not automatically guarantee the pod’s CPU cores and memory are local to that GPU unless topology-aware policies and resource managers are configured appropriately.
Node labels alone identify the right host, not necessarily the right socket inside a large server.
For tightly coupled workloads, configure CPU Manager/Topology Manager or an equivalent scheduler integration and test the actual cgroup/CPU set placement.
Virtualization and passthrough add another mapping layer
VMs with GPU passthrough or vGPU can hide some physical topology from the guest.
Hypervisor vCPU placement and guest NUMA topology should align with the passed-through GPU/NIC where performance depends on locality.
A guest can report healthy devices while the hypervisor schedules vCPUs on the remote socket and causes avoidable traffic across the host interconnect.
Inter-socket traffic should be monitored during performance tests
Use CPU performance counters, NUMA statistics, PCIe/NVLink/DCGM traffic, and application throughput to see whether a workload is crossing sockets unexpectedly.
Compare local and remote placement deliberately. Large improvements after pinning are strong evidence that topology, not the model kernel, was the bottleneck.
Keep the topology setting in the workload launcher or scheduler so the fix survives reboot and rescheduling.
BIOS and platform settings can affect locality
NUMA configuration, sub-NUMA clustering, PCIe bifurcation, memory interleaving, IOMMU, and power/performance settings can change how the operating system reports and uses topology.
Standardize BIOS/firmware across cluster nodes and include topology validation in hardware acceptance.
Two nominally identical servers with different BIOS settings can produce different CPU affinity and peer-access behavior.
NUMA optimization is successful when each workload uses its nearest resources by default
The mature platform maps GPU, NIC, NVMe, CPU, and memory locality; schedules or pins appropriately; validates transfers and collectives; and detects topology drift during burn-in.
NUMA should become an automated placement constraint, not a manual optimization remembered only after a benchmark underperforms.
IRQ and network-processing placement can undermine otherwise good GPU affinity. If RDMA NIC interrupts, softirq processing, or communication progress threads run on the remote socket, control overhead still crosses NUMA even when the NIC and GPU share a PCIe root. Review IRQ affinity and communication-library CPU binding during high-rate tests.
Memory interleaving can hide locality problems in benchmarks. Some BIOS or operating-system settings distribute allocations across sockets, which can make average throughput look acceptable while adding unpredictable latency. For performance-critical AI nodes, use a known NUMA policy and verify actual page placement instead of assuming the allocator chose local memory.
GPU process placement should be coupled with CPU allocation in Kubernetes or Slurm. A scheduler that gives one job four GPUs but only a small set of CPUs on the wrong socket can bottleneck preprocessing and communication. Resource templates should define CPU/memory expectations per GPU class and preserve locality where the scheduler supports topology-aware placement.
Multi-process service designs need extra care. One inference server can spawn workers per GPU while a parent process, networking thread, tokenizer, or scheduler remains on a different socket. Pin both data-plane workers and high-traffic service threads according to the topology, then verify with process-level CPU affinity and throughput measurements.
Topology can differ within one server family when GPUs or NICs occupy different slots. Do not hard-code “GPU0 uses NIC0” assumptions purely from device numbering. GPU numbering can change after firmware, driver, or hardware replacement; build mappings from PCI addresses/topology and stable device identity at startup.
NUMA tuning should be tested under oversubscription. A placement that performs well for one job can degrade when several jobs share host CPUs, memory controllers, and PCIe roots. Measure locality and throughput at the concurrency the node will actually host, especially for inference servers that run several model replicas per GPU node.
Storage, networking, and GPU affinity should be solved together. If a dataset comes through NIC0 near socket 0 but the assigned GPU lives under socket 1, moving only the CPU thread may not fix the path. Choose the GPU, NIC/storage path, CPU cores, and memory node as one locality bundle where possible.
A useful operational practice is to publish a per-node topology label or generated affinity map consumed by launch scripts. This turns NUMA from tribal knowledge into machine-readable placement intent and makes hardware changes visible when the generated map differs from the expected server profile.
CPU socket saturation can invalidate locality gains. Pinning every GPU worker to the “local” socket is not useful if that socket’s CPU cores or memory controllers are already saturated by data loaders and network interrupts. Balance locality with available CPU capacity, and benchmark a few sensible affinity layouts rather than assuming the closest node is always optimal under concurrency.
Huge pages and memory registration can be NUMA-sensitive as well. RDMA libraries, communication buffers, and storage stacks may allocate large pinned regions. Ensure those allocations occur on the intended node and that the platform does not fall back to remote pages because local memory is exhausted.
Topology-aware monitoring should flag unexpected SYS/NODE paths between GPUs and NICs after maintenance. A cable/riser or device replacement can move a NIC to another root complex without changing its Linux interface name. Automated topology checks protect against performance drift that ordinary link-health monitoring will never report.
NUMA policy should be included in reproducible job launchers rather than applied manually with one-off numactl commands. Slurm task affinity, Kubernetes CPU/Topology Manager policies, systemd units, or wrapper scripts should establish the intended CPU and memory binding automatically. Operationalizing affinity prevents performance fixes from disappearing the next time a job is rescheduled or a node reboots.
Compare topology with application phase. Training input pipelines may be most sensitive to CPU-memory locality, while distributed collectives are more sensitive to GPU-NIC locality and inference servers may be limited by tokenizer/network threads. There is no single perfect pinning layout for every workload, so platform defaults should be strong but overrideable with measured justification.
Schedulers and operators should verify placement with runtime evidence instead of trusting topology labels alone. CPU affinity, memory locality, PCIe attachment, and GPU selection should line up for the critical path; otherwise remote memory traffic can quietly erase expected acceleration.