RoCEv2 troubleshooting is hard because the failure often looks like an application problem: GPU jobs stall, checkpoint writes take longer, a storage node looks slow, or one rack suddenly drags down the rest of the pod. The network may show no classic packet loss on a normal interface graph, while the real issue is hiding in priority queues, ECN marking, PFC pause frames, NIC classification, or a backup flow accidentally sharing the lossless class.
This practical checklist is for small and mid-size AI clusters, data-center labs and private GPU pods where RoCEv2 is used for storage, distributed training or low-latency east-west traffic. It complements the AI Infrastructure hub, the Data Center Networking hub, and the broader Start Here networking topics page.