Skip to main content
Engineering write-up

GPUDirect RDMA over RoCE on bare-metal Kubernetes

What we changed, and why it mattered

A computer-vision team's 2-node training cluster was running gradient sync over 1GbE TCP. We specified RDMA-capable NICs and a DCB switch, then did the Kubernetes and NCCL configuration that makes RDMA actually get used.

Environment: 2-node bare-metal Kubernetes, 4 GPUs (RTX A5000/A5500), 100GbE RoCE.

Executive Summary

A computer-vision team trains object-detection and segmentation models on-prem for data-residency reasons. Their cluster: two workstations, four GPUs, Kubernetes on bare metal. Multi-node training was running gradient synchronization over the workstations' integrated 1GbE NICs — TCP, CPU in the data path, no RDMA capability at all.

The work had two halves. The hardware half: specify RDMA-capable 100GbE NICs and a DCB switch on a dedicated network (procurement and physical installation were the customer's). The software half — the part that usually goes wrong — make Kubernetes and NCCL actually use it: a second pod network via Multus, drivers and device plugin via NVIDIA Network Operator, PFC/ECN on the switch, and the NCCL GID index set so traffic doesn't silently fall back to TCP.

In clusters that already own RDMA hardware, that silent fallback is the failure we find most often in audits.

The Challenge

Client Context

The client, a computer-vision team, had built an on-premises ML platform to train object detection and image segmentation models. Privacy requirements and data residency regulations made cloud training impractical for their most sensitive workloads.

Their Infrastructure

  • 2 Dell Precision workstations: one with RTX A5000 (2 GPUs), one with RTX A5500 (2 GPUs)
  • Bare metal Kubernetes cluster (Ubuntu 24.04, K8s 1.29)
  • NVIDIA GPU Operator for device management
  • Calico CNI for pod networking
  • NFS for shared storage, local NVMe for scratch space

The Problem

Training jobs that used all 4 GPUs across both nodes were painfully slow. The GPUs spent much of each iteration idle, waiting on the network while gradients synchronized, with the CPU stuck in the data path for every byte transferred.

Root Cause Analysis

We identified three compounding issues:

1
No RDMA capability. The integrated NICs didn't support RDMA. Every gradient sync required: GPU memory → PCIe → System RAM → CPU (TCP/IP stack) → NIC → Wire → NIC → CPU → System RAM → PCIe → GPU memory. The CPU was in the critical path for every byte transferred.
2
No GPUDirect. Without GPUDirect RDMA, NCCL fell back to the Socket transport. Each AllReduce operation involved multiple memory copies and CPU intervention, adding latency per operation.
3
Inadequate bandwidth. Even ignoring latency, a 1GbE link cannot move gradient payloads of hundreds of megabytes fast enough for multi-node training to make sense.
Comparison diagram showing TCP/IP data path with multiple memory copies through CPU versus RDMA direct path bypassing CPU
TCP/IP Data Path vs RDMA Data Path

Two nodes with no RDMA path is a network problem, not a GPU problem. That meant RDMA.

Solution Architecture

Technology Selection: RoCE vs InfiniBand

For RDMA, there are two main options: InfiniBand and RoCE (RDMA over Converged Ethernet). We evaluated both:

FactorInfiniBand (HDR/NDR)RoCE v2 (100GbE)
Bandwidth200-400 Gb/s100 Gb/s
Latency~0.5-1 μs~1-2 μs
Switch cost$15-40K (IB switch)$3-8K (DCB Ethernet)
NIC cost~$1,500-3,000~$500-1,000
Expertise requiredSpecializedFamiliar to network teams

Decision: RoCE v2

For a 2-4 node deployment, RoCE offers 95% of InfiniBand's performance at 30% of the cost. The slight latency penalty (1-2μs vs 0.5-1μs) is negligible for gradient payloads measured in hundreds of megabytes.

Hardware Specification

ComponentSpecificationPurpose
NICNVIDIA ConnectX-6 Dx 100GbE (dual-port)RDMA-capable network interface
SwitchNVIDIA SN2201 or Dell S5248F-ONDCB-capable Ethernet with PFC/ECN
CablingDAC (Direct Attach Copper) or 100GbE QSFP28Node interconnect

Network Topology Design

We designed a physically separated network architecture with two distinct planes:

Physical network topology showing separate management and RDMA networks connecting GPU servers
Physical Network Topology

Management Network (existing)

Integrated NICs connected to the existing management switch. Handles Kubernetes control plane, SSH access, monitoring, and NFS traffic. No changes required.

RDMA Network (new)

ConnectX-6 Dx NICs connected to a dedicated DCB switch. Handles only GPU-to-GPU NCCL traffic. Flat L2 network with PFC/ECN enabled.

Why Flat L2 (No VLANs, No VXLAN)

  • No VLANs needed: With only 2 nodes on a dedicated switch, there's nothing to segment.
  • No VXLAN: Encapsulation overhead kills RDMA performance. VXLAN adds headers and processing that defeat the purpose of zero-copy transfers.
  • Simple PFC configuration: Priority Flow Control is easier to configure and debug on a flat network.

Kubernetes Integration

The Multi-Network Challenge

Kubernetes assumes a single network per pod. Our design requires pods to have two networks: the primary Calico network for Kubernetes services and a secondary RDMA network for NCCL traffic. This is where Multus CNI comes in.

Pod network architecture diagram showing eth0 for Calico and net1 for RDMA via Multus CNI
Pod Network Architecture with Multus CNI
eth0

Primary interface (Calico) for Kubernetes services, DNS, API server communication

net1

Secondary interface (RDMA) for NCCL collective operations

Component Stack

The complete Kubernetes stack for RDMA-enabled GPU training:

ComponentPurpose
CalicoPrimary CNI for pod networking
Multus CNIMeta-CNI for multiple network interfaces
NVIDIA Network OperatorRDMA drivers, device plugin, secondary networks
whereaboutsIPAM for secondary network
NVIDIA GPU OperatorGPU drivers, device plugin
KAI SchedulerGang scheduling for distributed jobs

SR-IOV vs Host-Device: A Critical Decision

ApproachHow It WorksProsCons
Host-DeviceEntire NIC moved into pod namespaceSimple, full performanceExclusive access - one job per NIC
SR-IOVVirtual Functions (VFs) carved from physical NICMultiple jobs share NICMore complex setup, ~5% overhead

Decision: Host-Device with Dual Ports

For this 2-node deployment running one distributed training job at a time, host-device provides the simplest path. The dual-port ConnectX-6 Dx gives us two RDMA resources per node, allowing two concurrent RDMA-enabled jobs if needed.

Distributed Training Setup

Framework: Kubeflow Training Operator

For running distributed PyTorch jobs on Kubernetes, we deployed the Kubeflow Training Operator. It provides PyTorchJob CRD for distributed PyTorch training, automatic worker discovery, rendezvous coordination, and integration with KAI Scheduler for gang scheduling support.

Distributed training stack diagram showing Kubeflow Training Operator, PyTorchJob, and KAI Scheduler integration
Distributed Training Stack

NCCL Configuration for RoCE

NCCL (NVIDIA Collective Communication Library) must be explicitly configured to use the RDMA interface. Key environment variables include enabling the InfiniBand/RoCE path, specifying the ConnectX-6 device name, and setting the correct RoCEv2 GID index.

What changed

Gradient synchronization moved off the CPU and onto the RDMA path; AllReduce stopped being the step the GPUs waited on. Multi-node jobs that had been network-bound became GPU-bound, which is the state you want. The team's multi-day fine-tuning runs became same-day runs, and multi-node training went from something they avoided to the default.

The configuration detail that mattered most was the smallest one: the NCCL GID index. Get it wrong and everything runs — over TCP, silently, at a fraction of the speed.

When to Use This Architecture

This solution is appropriate when:

  • You're running multi-node distributed training (DDP, FSDP, DeepSpeed)
  • Gradient payloads exceed 100MB (most modern models)
  • GPU utilization during training is under 60%
  • You control the hardware (on-prem, colo, bare metal cloud)

This solution is overkill when:

  • Training fits on a single node
  • You're doing inference only
  • You're on shared cloud infrastructure without RDMA support