A slow AllReduce is not automatically a slow switch. Before buying faster NICs or copying NCCL tuning variables, isolate whether the problem follows a GPU pair, a server, a rail, a rack boundary, or the application. This guide provides a reusable troubleshooting matrix, a staged test plan and a worked bandwidth example for an AI training cluster.
Scope: NVIDIA Collective Communications Library (NCCL), GPU servers using InfiniBand or RDMA over Converged Ethernet (RoCE), and the boundary between local GPU communication and scale-out networking. All designs, timing inputs and acceptance policies below are hypothetical. Commands are unexecuted examples for an authorized maintenance allocation—not reported benchmark runs. Documentation was checked on September 11, 2026; match settings to your installed NCCL, CUDA, driver and platform versions rather than assuming rolling documentation describes your deployment.