AI cluster network telemetry should answer one operational question quickly: is the training or inference job slow because of the application, storage, GPU scheduling, or the network fabric? In RoCEv2 and high-throughput Ethernet fabrics, waiting for a hard failure is too late. By the time users report idle GPUs, the useful evidence may already be spread across switch counters, NIC statistics, job logs and monitoring gaps.
This practical checklist focuses on signals that catch congestion, packet loss and unstable paths before they become a full incident. It is vendor-neutral on purpose: the exact commands change, but the telemetry model is reusable for leaf/spine data centers, AI infrastructure pods and lab environments.