In 20 years, you will be more dissapointed by what you didn't do than by what you did.

Showing posts with label InfiniBand. Show all posts
Showing posts with label InfiniBand. Show all posts

NVIDIA InfiniBand for AI Clusters: NDR vs XDR, Rail Design and Acceptance Tests

A GPU cluster can have every cable connected and still be unready for training. The useful engineering question is not simply “NDR or XDR?” It is whether the port map, subnet management, routing policy, GPU-to-NIC alignment and failure behavior describe the same system.

This focused AI infrastructure installment develops a hypothetical 64-server, 512-GPU InfiniBand design, then turns it into a commissioning workflow. Calculations below are theoretical capacity accounting, executed with Python—not benchmark results. The example assumes eight GPUs and eight independent 400Gb/s network ports per server; it is not a claim that every eight-GPU server has that topology.

Read More ->>

Popular Posts