NVIDIA NCA-AIIO: NVLink and GPU Interconnects

NVIDIA NVLink is a high-bandwidth GPU interconnect designed for tightly coupled GPU memory and collective communication. Current GB300 NVL72 reference architecture uses fifth-generation NVLink, delivering up to 1,800 GB/s bidirectional bandwidth per GPU and connecting 72 Blackwell Ultra GPUs through NVSwitch into one rack-scale NVLink domain. NVLink solves a different problem from PCIe, InfiniBand, or Spectrum-X Ethernet: it is the local/rack-scale GPU fabric for extremely high-bandwidth GPU-to-GPU communication, while network fabrics connect nodes/racks and external resources.

Within NVIDIA AI Infrastructure, interconnect design should follow the communication hierarchy: GPU memory local to one accelerator, NVLink/NVSwitch inside a domain, PCIe/CPU paths where unavoidable, then GPUDirect RDMA over InfiniBand or Spectrum-X between systems.

NVLink bypasses the limits of PCIe peer traffic

PCIe remains essential for host/device and many peripheral connections, but it is not optimized for the bandwidth/latency of tightly synchronized multi-GPU training and inference.

NVLink provides direct GPU connectivity with substantially higher aggregate bandwidth on supported systems.

Topology-aware libraries choose NVLink paths when available rather than staging every transfer through CPU memory.

NVSwitch turns point links into a fabric

NVSwitch provides switching across multiple GPU NVLink ports so many GPUs can communicate through a consistent high-bandwidth domain.

In GB300 NVL72, current NVIDIA architecture uses nine fifth-generation NVLink switch trays, each with two NVSwitch ASICs.

This gives all 72 rack GPUs access to one NVLink domain instead of independent small GPU pairs.

Fifth-generation NVLink changes rack-scale model placement

Current NVIDIA GB300 reference architecture exposes 900 GB/s one-direction / 1,800 GB/s bidirectional NVLink bandwidth per GPU.

Large models can be partitioned across GPUs with much more local communication bandwidth than ordinary network fabrics provide.

Use this capacity for tensor/model parallelism, KV/state exchange and collective patterns that would otherwise be network bound.

NVLink does not replace the scale-out network

An NVL72 rack still requires Spectrum-X or InfiniBand/Ethernet for communication to other racks, storage, services and management.

InfiniBand Fabric Tuning covers one scale-out option; NVIDIA Spectrum-X Ethernet covers the Ethernet option.

Keep scale-up and scale-out domains explicit in topology diagrams and performance troubleshooting.

NCCL should see the true topology

NCCL Collective Performance depends on correct discovery of NVLink/NVSwitch, PCIe, NUMA and network paths.

Containers/VMs should expose enough topology for NCCL to choose efficient paths.

Forced transport settings can accidentally route communication away from an available NVLink path.

NVLink health is part of GPU health

A GPU can execute compute correctly while one interconnect link is degraded.

Use DCGM and platform telemetry to monitor link state and relevant error counters.

A single bad link can increase collective tail latency, which slows every rank in a synchronized job.

GPU placement should respect NVLink domains

Schedulers should keep tightly coupled jobs inside the smallest NVLink/NVSwitch domain that satisfies GPU count before crossing a slower scale-out boundary.

When a job needs more GPUs than one domain, topology-aware placement should minimize cross-domain traffic.

This matters for predictable performance and network capacity planning.

NVLink partitioning creates isolation boundaries

Current NVL72 terminology includes NVLink domains, blocks and partitions managed through fabric-management services.

Partitions can isolate groups of nodes/GPU memory access inside a shared rack-scale fabric.

Coordinate partition state with scheduler/job allocation so infrastructure isolation and workload allocation describe the same domain.

Power and cooling failures can look like interconnect performance problems

Dense rack-scale GPUs can throttle or fail under thermal/power constraints, reducing collective performance even when NVLink topology is correct.

Power and Cooling for AI Racks should be investigated alongside link counters when rack-wide throughput changes with environmental load.

Interconnect tuning cannot compensate for GPU throttling.

Benchmark scale-up separately from scale-out

Run GPU peer/NVLink microbenchmarks inside the NVLink domain, then network/RDMA tests across nodes/racks, then end-to-end collectives.

This isolates whether the bottleneck is local interconnect or scale-out fabric.

Store golden baselines per hardware generation because NVLink generations and topology differ significantly.

NVLink succeeds when schedulers and libraries treat topology as a first-class resource

The mature design maps NVLink/NVSwitch domains, monitors link health, places jobs topology-aware, uses NCCL automatic transport selection, and separates scale-up from scale-out troubleshooting.

GPU count alone is not a useful capacity metric; the interconnect topology determines how those GPUs behave as one system.

Topology documentation should identify every GPU’s NVLink peers, NVSwitch path, PCIe root, CPU/NUMA relationship and scale-out NIC. A job scheduler and performance engineer need the same map. Without it, an apparently uniform 8- or 72-GPU pool can contain paths with very different local and remote bandwidth.

NVLink generation matters. Bandwidth and switch-domain capabilities differ substantially across Volta, Ampere, Hopper and Blackwell systems. Keep performance targets by hardware generation and do not use an NVL72 fifth-generation NVLink expectation for an older HGX node. Capacity planning should compare like-for-like systems.

GPU reset or maintenance can affect the wider NVLink domain depending on platform and fabric management. In rack-scale systems, coordinate service/drain procedures with NVLink partition/block ownership so one repair does not interrupt unrelated jobs unnecessarily. Maintenance tooling should understand the fabric domain, not just the PCIe device.

NVSwitch telemetry should be monitored alongside GPU link counters. Switch ASIC or link-tray faults can degrade several GPU paths simultaneously. Correlate common failures across GPUs to identify a shared NVSwitch or tray problem instead of replacing several healthy GPUs that all report communication errors.

Collective performance should be compared inside one NVLink domain before scaling across domains. If an AllReduce is slow on a single rack, fix NVLink/NCCL/topology first. If single-domain performance is healthy but multi-rack jobs are slow, investigate scale-out fabric and placement. This layered benchmark prevents expensive network changes for a local problem.

Model-parallel frameworks should be topology aware. Tensor-parallel groups benefit from the highest-bandwidth local links, while data-parallel replicas can often tolerate a slower boundary. Place communication-heavy dimensions inside NVLink domains and use scale-out networking for less chatty synchronization where the framework supports that mapping.

Power-management states can change observed NVLink performance. Idle/downclocked GPUs may show lower benchmark numbers until warmed, while thermal/power throttling can degrade sustained throughput. Use controlled benchmark duration and DCGM clock/throttle telemetry when establishing a golden baseline.

Virtualization and partitioning require supported interconnect behavior. vGPU/MIG/VM configurations may expose a different peer-access model from bare metal. Validate P2P/NVLink visibility in the exact virtualization mode rather than assuming the physical topology is fully visible to each guest or container.

Fabric-management APIs should be controlled like other infrastructure control planes. NVLink domain/partition changes can alter which GPUs can access each other’s memory and which jobs are isolated. Restrict administrative access, log changes, and coordinate them with scheduler state.

Interconnect capacity planning should forecast communication per model architecture, not only GPU count. Mixture-of-experts, tensor parallelism and long-context inference can create very different traffic patterns from independent batch inference. The right NVLink domain size depends on the model’s communication structure.

NVLink bandwidth should be considered alongside GPU memory capacity. Very large models may require distributing parameters or KV cache across several GPUs even when compute load is moderate. High local interconnect bandwidth can make such scale-up feasible, while a lower-bandwidth topology may force different parallelization or quantization choices.

Topology-aware admission control can prevent pathological placements. If a job asks for 8 tightly coupled GPUs, the scheduler should prefer 8 within one NVSwitch domain rather than 4+4 across domains unless the job explicitly supports the slower boundary. This turns topology into a placement constraint rather than a post-hoc performance explanation.

Link health should be checked after physical service. Replacing a GPU tray, switch tray or cable can restore basic connectivity while leaving one path at reduced capability. Run the golden NVLink benchmark and compare DCGM topology/link state before returning the rack to full production.

Scale-up network failures can cascade into application timeouts instead of obvious hardware errors. Collective libraries may retry or reroute until synchronization stalls. Correlate NCCL RAS/debug, DCGM link counters and workload timestamps to identify interconnect faults before application teams rewrite communication code.

Interconnect utilization can inform model-parallel tuning. If NVLink is saturated while GPU compute has headroom, reduce communication or adjust parallel dimensions; if NVLink is lightly used but GPUs idle on remote synchronization, the bottleneck may be scale-out network or rank imbalance instead.

Document supported peer-access behavior for every node type in a heterogeneous cluster. A workload moved from one generation to another may see a different NVLink/NVSwitch topology despite the same GPU count. Scheduler resource classes should preserve those distinctions.

Interconnect-aware cost modeling should consider completed work per rack, not just GPU purchase price. A denser NVLink domain can reduce communication overhead and improve expensive accelerator utilization, but it also increases facility and rack-scale dependency. Compare job throughput and failure-domain economics together.

Keep topology baselines and link-health thresholds under version control with the hardware generation.

Keep scheduler topology labels synchronized with the physical NVLink domain after maintenance or hardware replacement. A stale label can place a tightly coupled job across the wrong boundary even though the rack itself was repaired correctly.

Topology-aware scheduling should be validated after failures and maintenance as well as on a pristine cluster. GPU locality can change when nodes, links, or partitions are unavailable, and the scheduler must continue placing communication-heavy workloads on paths that match their performance assumptions.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!