A computer-vision team trains object-detection and segmentation models on-prem for data-residency reasons. Their cluster: two workstations, four GPUs, Kubernetes on bare metal. Multi-node training was running gradient synchronization over the workstations' integrated 1GbE NICs — TCP, CPU in the data path, no RDMA capability at all.
The work had two halves. The hardware half: specify RDMA-capable 100GbE NICs and a DCB switch on a dedicated network (procurement and physical installation were the customer's). The software half — the part that usually goes wrong — make Kubernetes and NCCL actually use it: a second pod network via Multus, drivers and device plugin via NVIDIA Network Operator, PFC/ECN on the switch, and the NCCL GID index set so traffic doesn't silently fall back to TCP.
In clusters that already own RDMA hardware, that silent fallback is the failure we find most often in audits.