Skip to main content

GPU Networking & RDMA

The network between your GPUs is the single biggest performance lever in distributed training. A misconfigured switch port or missing PFC config silently kills throughput for the entire cluster. We design and implement RDMA networks that run at wire rate.

Lossless by design

InfiniBand or RoCE v2 fabrics configured for real workloads: PFC, ECN/DCQCN, QoS and MTU tuned across NICs, switches and hosts.

Zero-copy across nodes

GPUDirect RDMA verified across the path, so inter-node GPU traffic moves off the CPU entirely.

Forensic when it's broken

We bring perftest, nccl-tests and switch counter experience to find why the fabric underperforms, then fix it in place.

What we do

Design, build and repair the GPU fabric.

From switch topology down to the driver on each NIC, we make the RDMA path work and stay working.

01
InfiniBand fabric
Quantum switch deployment, subnet manager configuration, adaptive routing, fat-tree and dragonfly topology design, partition keys.
02
RoCE v2 fabric
Lossless Ethernet with PFC, ECN/DCQCN tuning, leaf-spine design, ECMP multi-path, jumbo frames and DSCP trust.
03
GPUDirect RDMA
Zero-copy GPU-to-GPU transfers bypassing the CPU, peer memory module setup, GDR copy validation and firmware tuning.
04
Switch configuration
Spectrum-X and Quantum switch deployment, port speed validation, error counter monitoring, QoS policies and MTU config.
05
Kubernetes networking
Network Operator (NicClusterPolicy, RDMA device plugin), Multus secondary networks, SR-IOV, and dual-network designs that separate management from RDMA training traffic. We contributed the global config feature upstream.
How we work

Scope, design, build, hand off.

01

Assess

Audit link speeds, error counters, PFC/ECN and PCIe topology.

02

Design

Fabric topology, oversubscription, QoS and traffic separation.

03

Implement

Configure switches, NICs, RDMA and GPUDirect. Validate across the path.

04

Transfer

Network monitoring dashboards, runbooks and documentation.

Our engineers contribute upstream to the projects this layer runs on: KAI Scheduler, Network Operator, DOCA driver build, ipoib-cni. See the contributions →
Technologies we work with
InfiniBandRoCE v2GPUDirect RDMAConnectX-6/7Spectrum-XQuantumNCCLNetwork OperatorMultusSR-IOVMACVLAN
FAQ
What is RDMA and why does it matter for GPU training?
RDMA lets NICs read and write remote memory directly, bypassing the CPU and kernel. Combined with GPUDirect RDMA, it enables zero-copy GPU-to-GPU transfers across nodes, giving far higher effective bandwidth and an order-of-magnitude lower latency than TCP on the same hardware.
Should I use InfiniBand or RoCE?
Both deliver RDMA performance. InfiniBand is a purpose-built lossless fabric standard in DGX SuperPOD deployments. RoCE v2 runs RDMA over Ethernet: cheaper, more flexible, and the right choice for most cloud, colo and bare-metal clusters when PFC and ECN are configured correctly.
Do I need PFC and ECN for RoCE?
Yes, if you want lossless RoCE v2. PFC prevents packet drops during microbursts and ECN signals congestion before buffers overflow. Without these configured across NICs, switches and host settings, RoCE falls over under load and NCCL silently underperforms.
What is GPUDirect RDMA?
GPUDirect RDMA lets the NIC DMA directly to and from GPU memory without an intermediate CPU copy. It requires matched driver support, peer-memory modules and PCIe affinity between GPU and NIC. When enabled, inter-node GPU communication drops to single-digit microseconds.
Can you fix existing GPU networking problems?
Yes. A lot of our work is forensic: NCCL falling back to TCP, PFC dropping packets under load, GPU-to-NIC PCIe affinity mismatches, wrong NCCL_IB_HCA, incorrect DSCP marking. We bring perftest, nccl-tests and switch counter experience to find and fix these without replacing hardware.
Do I need two NICs per GPU node?
For production distributed training, yes. One NIC for Kubernetes management (pod CNI, API traffic, metrics) and one or more RDMA-capable NICs dedicated to NCCL and training traffic via Multus secondary networks. Single-NIC works for proofs of concept but degrades at scale.
Network holding back your GPUs?

We'll profile your fabric, find the bottleneck and fix it.