In 20 years, you will be more dissapointed by what you didn't do than by what you did.

GPUDirect RDMA Slow or Not Working? GPU-to-NIC Troubleshooting Matrix

A GPU cluster can have active network ports and still fail to move GPU-resident data efficiently. The useful search question is not simply “is InfiniBand working?” It is: does the intended GPU-to-NIC path work, and what changes when the buffer moves from host memory to GPU memory? NVIDIA recommends that host-versus-GPU comparison when troubleshooting GPUDirect RDMA.[5]

This guide provides a staged test matrix, a hypothetical 256-GPU coverage worksheet and an evidence-based acceptance checklist. It targets the boundary between server topology, GPU-memory registration and network transport—not general collective tuning. All hardware layouts and planning thresholds below are illustrative. Commands are unexecuted examples; only the sizing worksheet was executed for this article. No GPU benchmark results are claimed.

This is a focused companion to the NCCL Slow AllReduce troubleshooting guide: here the reusable asset is a GPU-memory registration and endpoint-coverage plan, rather than a collective-bandwidth model. Public documentation was checked on October 9, 2026. These are rolling documentation pages; match every option and support condition to your installed versions, not the page's current version banner.

1. Separate three paths before changing settings

GPU-to-GPU peer access inside a server, GPU-to-NIC RDMA and inter-node fabric connectivity are separate diagnostic questions. NVIDIA describes GPU peer-memory access, GPU-to-NIC communication and topology discovery separately because configuration problems can affect them differently.[4]

GPUDirect RDMA enables compatible third-party devices to access GPU memory through the supported mapping and driver mechanisms. NVIDIA describes nvidia-peermem as providing supported InfiniBand adapters direct peer-to-peer read/write access to GPU video memory without copying through host memory.[2]

That is not proof that an application selected this path. Treat a passing host-memory RDMA test as evidence about that test, not evidence that GPU-memory registration or application transport selection succeeded. Similarly, an intra-node GPU test does not exercise the leaf/spine network. Our recommended investigation keeps these claims distinct throughout the incident record.

Begin with one failing pair of nodes and one known-good comparison pair. Use the same software image, message sizes and direction. Record whether the failure follows a GPU position, NIC, node, rack, container or job definition. This is an experimental workflow, not a claim that any particular component is faulty.

Four diagnostic stages from host memory to application tests
Source: original Network freak diagram; illustrative diagnostic workflow and hypothetical planning inputs, not measured results.

2. Freeze the test envelope and inventory

Before benchmarking, save the actual GPU model and form factor, server platform, NIC model and port mode, firmware, operating-system kernel, GPU driver flavor/version, CUDA runtime, NCCL version, network plugin and perftest build. Do not substitute a purchase-order description for the discovered inventory. Document the host and container separately.

For a reproducible incident bundle, capture the following read-only commands on both endpoints. Availability and permissions vary by distribution; these are examples to adapt, not an installation script:

nvidia-smi
nvidia-smi topo -m
lspci -tv
lspci -vv
ibstat
ibstatus
rdma link show
rdma statistic
ib_write_bw --help
ulimit -l

NVIDIA recommends checking active port state, link layer and expected link rate before bandwidth testing.[5] NCCL also relies on /sys for PCI topology discovery, so a container or virtual machine exposing an inaccurate topology can produce suboptimal selection.[4]

Create a mapping with one row per intended GPU/NIC association: physical GPU identifier, process-visible GPU index, PCI bus address, RDMA device, port, NUMA locality, rail and switch attachment. Treat process indices as scoped names: confirm them inside the actual job environment before choosing test devices. Keep production identifiers in a private incident attachment rather than a public report.

A topology display is a hypothesis generator, not a performance certificate. Use it to select local and remote comparison paths; then test those paths. Do not declare the platform healthy because the diagram resembles a reference architecture.

3. Choose the supported GPU-memory registration route

There are two relevant routes to check, not one universal module test. The NCCL documentation describes nvidia-peermem and also a DMA-BUF route using supported recent kernels and NVIDIA's open-source GPU driver; when DMA-BUF is used, the peer-memory module is unnecessary.[4]

Therefore, an absent nvidia-peermem module is not by itself a fault. First identify which route the installed software stack is intended to use. For the peer-memory route, check the module and RDMA integration against the platform support instructions. NVIDIA's GPUDirect guide notes that competing legacy nv_peer_mem and nvidia-peermem use must be resolved, and that installation order can matter when the GPU driver needs RDMA APIs supplied by MLNX_OFED.[2]

For DMA-BUF, validate the actual kernel, GPU driver and perftest build rather than assuming that a recent package name is sufficient. The perftest README documents CUDA build support and the combination of --use_cuda with --use_cuda_dmabuf for the DMA-BUF path.[3]

Recommended decision record:

Question Evidence to retain Stop condition
Which registration route is intended? Platform support document and selected software bill of materials No supported combination identified
Does the binary support the requested GPU options? Local help output and build/version details Missing options or incompatible builds
Does GPU-memory registration succeed? Complete server and client logs Allocation/registration error
Is the same route used inside the job? Runtime inventory and application diagnostics Host passes but container/job differs

Do not install multiple alternative driver stacks during one diagnostic session. Our recommendation is to preserve the current image, change one supported component at a time and rerun the identical pair test. Otherwise, a successful result cannot tell you which change mattered.

4. Run host-memory and GPU-memory tests separately

The perftest project describes its tools as synthetic microbenchmarks, not emulations of real application traffic, and requires matching options on server and client.[3] Use that distinction in the acceptance report: a pair test establishes a bounded transport result, not training throughput.

For a first InfiniBand pair test, choose an existing active HCA and a reachable peer bootstrap address. The example device name is a placeholder. Use the corresponding device name on each host; hardware enumeration need not match. The shell variable must be assigned to the intended peer address before running the client.

# Example only: server, host-memory baseline
ib_write_bw -d mlx5_0 -a

# Example only: client, matching host-memory baseline
ib_write_bw -d mlx5_0 -a "$PEER_BOOTSTRAP_ADDRESS"

Repeat with GPU memory on both endpoints, using the correct process-visible GPU index on each:

# Example only: server, GPU-memory path
ib_write_bw -d mlx5_0 --use_cuda=0 -a

# Example only: client, GPU-memory path
ib_write_bw -d mlx5_0 --use_cuda=0 -a "$PEER_BOOTSTRAP_ADDRESS"

When the supported stack uses DMA-BUF, add --use_cuda_dmabuf on both sides. These options are documented by perftest and NVIDIA, but their presence depends on the installed build.[3][5] On RoCE, additionally validate the intended network, addressing and applicable GID selection using your version-specific procedure; these minimal InfiniBand examples are not a universal RoCE recipe.

Repeat with endpoint roles reversed, and retain the full size sweep rather than only its highest number. Record units exactly as the tool prints them. Do not compare one-direction traffic to an aggregate bidirectional number. Keep CPU affinity, message sizes, queue settings and test duration fixed when comparing memory paths.

If the GPU-memory invocation fails before transferring data, a throughput chart is premature. Resolve the registration or device-selection failure first. If both memory paths work, compare the shapes of the curves, repeatability and counter changes. Our recommended next experiment is the intended local GPU/NIC pairing against a supported but less-local pairing, with every other variable held constant.

5. Troubleshooting matrix: what to investigate next

The following matrix is a proposed triage asset, not a list of observed faults. Its branches follow NVIDIA's separation of registration, topology and fabric checks.[4][5]

Observation Next controlled check Avoid this conclusion
Host and GPU tests both fail Port state, bootstrap reachability, fabric configuration and matching test options “The GPU driver must be broken”
Host passes; GPU registration fails Supported registration route, GPU visibility, build flags and effective memory limits “More switch bandwidth will fix it”
Host performs consistently; GPU varies by slot GPU/NIC mapping, PCI hierarchy and supported locality comparison “Every active NIC is equivalent”
Bare metal passes; container fails Device exposure, runtime limits, libraries and /sys topology “The whole fabric is congested”
Pair tests pass; concurrent pairs degrade Shared server resources, rail utilization and fabric counters “Single-pair speed proves full-cluster capacity”
Verbs tests pass; NCCL initialization hangs Bootstrap interfaces, TCP connectivity, runtime and transport logs “RDMA success validates process startup”
One rail repeatedly underperforms Rail-specific topology and port/counter evidence “Disable the rail permanently”

NVIDIA notes that NCCL can select an interface that is up but not actually reachable, and that NCCL uses TCP ports to exchange connection information.[5] This is why bootstrap connectivity belongs in the checklist even when the payload transport is RDMA.

Memory-registration failures also deserve a scoped check of the job's effective locked-memory limit. NVIDIA documents pinned-memory limits as a cause of registration-related errors and emphasizes that updated limits must reach newly launched jobs.[5] Do not assume that editing a login configuration changes a running scheduler service or existing container. Verify the limit from the execution context that fails.

6. ACS, IOMMU and isolation: do not apply a blanket fix

PCI Access Control Services (ACS), IOMMU configuration and platform topology can affect direct PCI traffic. NVIDIA warns that ACS can redirect peer traffic toward the CPU root complex and discusses different requirements for bare-metal and virtual-machine environments.[4]

This is a supportability and isolation decision, not a copy-and-paste tuning trick. Do not run a script that clears ACS across every bridge on a shared host. Do not disable IOMMU globally because one bandwidth measurement is low. Obtain a platform-specific supported configuration and assess device assignment and tenant-isolation requirements before changing firmware or kernel settings.

The documentation also distinguishes Linux bare-metal CUDA PCIe GPU-to-GPU peer access from virtualized configurations.[4] Avoid turning that guidance into a universal assertion about every GPUDirect RDMA platform. Record the specific transfer path, whether the system is bare metal or virtualized, the relevant vendor guidance and the rollback procedure.

Our recommended change gate requires a drained test node, saved firmware/kernel settings, out-of-band access, an approved maintenance window and an explicit post-change isolation review. If the supported high-performance mode conflicts with the required tenant boundary, select a different deployment architecture rather than silently weakening isolation.

7. A 256-GPU coverage and bandwidth worksheet

Hypothetical design: 32 servers, eight GPUs and eight single-port 400 Gb/s NIC associations per server, arranged across eight rails. These are planning inputs, not specifications of a named product. Assume the selected server can expose these resources in a supported configuration; actual PCIe and platform capacity still require verification.

For the first coverage pass, pair the servers into 16 disjoint pairs. Exercise eight intended GPU/NIC associations per pair, reverse direction and repeat three times. This gives 768 directed pair-test executions. At an assumed 30 seconds per execution, serial runtime is 6.4 hours. Running all 16 disjoint server pairs concurrently reduces the idealized active-test time to 24 minutes, excluding setup, warm-up, retries and evidence collection.

Worksheet quantity Calculated result Meaning
GPUs 256 Inventory coverage target
Intended GPU/NIC associations 256 One association per GPU in this example
Disjoint server pairs 16 Each server participates once per round
Pair-test executions 768 16 pairs × 8 rails × 2 directions × 3 repeats
Nominal per-port rate 50 GB/s 400 Gb/s divided by eight; theoretical
Nominal per-server sum 400 GB/s Eight independent port rates; theoretical
Nominal cluster endpoint sum 12.8 TB/s 32 server sums; not fabric bisection bandwidth
Idealized concurrent active-test time 24 minutes 768 × 30 seconds / 16 workers

The endpoint sum is deliberately labeled: it is neither a measured collective result nor a guarantee that the servers, fabric or workload can sustain it. It also is not the throughput of this disjoint-pair schedule, in which only a subset of endpoint ports is active at a time. Do not add it to a bidirectional NVLink number or use it as a training-throughput forecast.

Executable planning calculator—standard-library Python, with decimal bandwidth units:

servers, gpus_per_server, rails = 32, 8, 8
rate_gbps, repeats, seconds = 400, 3, 30
assert servers % 2 == 0
pairs = servers // 2
runs = pairs * rails * 2 * repeats
print("GPUs:", servers * gpus_per_server)
print("Directed executions:", runs)
print("Serial hours:", runs * seconds / 3600)
print("Ideal parallel minutes:", runs * seconds / pairs / 60)
print("Port GB/s:", rate_gbps / 8)
print("Server GB/s:", rails * rate_gbps / 8)
print("Endpoint sum TB/s:", servers * rails * rate_gbps / 8 / 1000)

This schedule validates every intended endpoint association against one peer, not every possible server pair or switch path. Add a second pairing round across different racks, then a concurrency phase that exercises all intended rails. Keep those phases separate: endpoint coverage, path diversity and congestion exposure answer different questions. A complete hardware census is not complete topology coverage.

Endpoint coverage schedule for a hypothetical 256 GPU cluster
Source: original Network freak diagram; illustrative diagnostic workflow and hypothetical planning inputs, not measured results.

8. Failure domains and alternative deployment choices

Classify failures at four scopes: GPU/NIC association, whole server, rail and application environment. A defect isolated to one association calls for different containment than a rail-wide issue. Our suggested operational response is to quarantine only the smallest proven affected allocation unit, preserve evidence and retest before broad replacement.

The eight-NIC example is not a purchasing recommendation. A platform with fewer NICs may be appropriate when the workload, locality and budget support it; the acceptance plan must then include simultaneous traffic from GPUs sharing each resource. Conversely, adding NICs is not evidence that a shared PCI path or application mapping can use them efficiently.

For network choice, the peer-memory mechanism described by NVIDIA can support both InfiniBand and RoCE on compatible adapters.[2] Choose the fabric using your operational requirements, supported platform and workload tests—not the presence of the word GPUDirect. For the Ethernet branch, use the companion RoCE PFC buffer sizing worksheet rather than transplanting InfiniBand management checks into an Ethernet incident.

If connectivity fails only between partition members, continue with the InfiniBand P_Key troubleshooting matrix. Keep partition policy diagnosis separate from memory-registration diagnosis.

9. Deployment and acceptance checklist

Use this proposed checklist as a handover artifact. Set site-specific numerical targets before testing; there is no universal percentage of nominal line rate that proves GPUDirect healthy.

  • Inventory: capture supported firmware, driver, kernel, runtime and test-tool combinations; retain host and job views.
  • Mapping: account for every intended GPU/NIC association and its rail; verify device visibility within the real allocation.
  • Connectivity: confirm active links, expected mode/rate, correct fabric policy and working process bootstrap.
  • Registration: document the intended peer-memory or DMA-BUF route; prove GPU-memory tests start and complete.
  • Pair coverage: test host and GPU buffers, both endpoint roles and repeatability; archive complete logs.
  • Concurrency: repeat with the intended number of active rails and jobs, watching counters during the same time window.
  • Application: test representative collective message sizes and the actual framework after lower-layer checks pass.
  • Isolation: approve any firmware or kernel change with security/platform owners and verify rollback.
  • Recovery: repeat the approved test set after reboot and after a controlled supported failure/recovery exercise.
  • Handover: attach actual measurements, deviations, owners and retest dates; never substitute this theoretical worksheet for results.

For counter review, NVIDIA recommends RDMA statistics and fabric-specific diagnostics; on RoCE it also calls out PFC, ECN, CNP and queue-drop observations when investigating congestion.[5] Prefer before/after deltas tied to the test window rather than treating every historical nonzero counter as a new fault.

Practical takeaways

Prove the host-memory path, then the GPU-memory path, then the concurrent application path. Record which registration mechanism is supported instead of diagnosing by module presence alone. Use the 256-GPU worksheet to budget coverage and test time, not to predict training speed. Keep isolation-sensitive changes behind an explicit approval gate.

Continue with the 256-GPU cluster acceptance plan, or browse the AI infrastructure and data-center networking hubs. This article's durable assets are the memory-path comparison workflow, symptom matrix, executable coverage worksheet and deployment checklist.

Sources

Comments

0 Responses to "GPUDirect RDMA Slow or Not Working? GPU-to-NIC Troubleshooting Matrix"

Post a Comment

Popular Posts