AI cluster networking often fails in ordinary places: an underlay route is missing, an EVPN route target is wrong, a VTEP is silent, or storage traffic accidentally shares a failure domain with GPU traffic. The hardware may be expensive, but the most reliable design pattern is still a boring, testable leaf-spine fabric with clear separation between underlay, overlay and tenant policy.
This practical checklist is written for small and medium AI labs, private GPU pods and data-center teams that want a repeatable BGP EVPN/VXLAN design without turning every troubleshooting session into a vendor-specific archaeology project.