Skip to main content

Distributed Training Optimization

Multi-node training that barely scales past a single node. GPUs sitting idle during AllReduce. NCCL timeouts killing overnight runs. We fix the network and configuration layer that causes all of this.

The network is the lever

Most underperforming clusters are misconfigured, not under-provisioned. We recover throughput on the NICs and switches you already own.

Verified, not guessed

We profile with nccl-tests and perftest, then tune against measured AllReduce-to-compute ratios rather than defaults.

We do the implementation

NCCL algorithm selection, RDMA fabric config and topology-aware placement, applied and validated on your cluster.

What we do

The path your gradients travel.

We tune every hop from GPU memory to the wire and back, so collective communication stops stalling the job.

01
NCCL tuning
Algorithm selection (Ring, Tree, CollnetDirect), protocol tuning, buffer sizing and thread configuration for your specific topology.
02
RDMA / RoCE configuration
PFC, ECN/DCQCN, GID indexes, traffic class and DSCP marking, with lossless validation across NICs, switches and hosts.
03
InfiniBand optimization
Subnet manager config, adaptive routing, partition keys and rail-optimized topologies.
04
GPUDirect RDMA setup
Zero-copy GPU-to-GPU transfers, peer memory modules and GDR copy validation.
05
Topology and profiling
NVLink/NVSwitch intra-node routing, PCIe affinity and NUMA-aware placement, plus perftest and nccl-tests to find the real bottleneck.
How we work

Scope, design, build, hand off.

01

Assess

Profile your network fabric, NCCL config and GPU topology.

02

Diagnose

Find the real bottleneck. It is usually the network.

03

Implement

Tune NCCL, configure RDMA and fix switch configs.

04

Transfer

Document everything so your team operates independently.

Our engineers contribute upstream to the projects this layer runs on: KAI Scheduler, Network Operator, DOCA driver build, ipoib-cni. See the contributions →
Technologies we work with
NCCLInfiniBandRoCE v2GPUDirect RDMAConnectX-6/7PyTorch DDPFSDPDeepSpeedMegatron-LMH100A100GH200
FAQ
What is distributed training optimization?
It is the practice of tuning GPU networking, NCCL and collective-communication paths so multi-node training scales near-linearly with node count. That means RDMA/RoCE configuration, GPUDirect RDMA, NCCL algorithm tuning and topology-aware process placement to eliminate network-induced GPU idle time.
How do I know if my multi-node training is network-bound?
If scaling from 8 to 64 GPUs delivers far less throughput than the added GPUs should, or GPUs sit idle during AllReduce, the network is the bottleneck. Profiling with nccl-tests and per-iteration timing of AllReduce versus compute will confirm it.
What is the difference between NCCL over TCP and NCCL over RDMA?
NCCL over TCP goes through the kernel networking stack. NCCL over RDMA (RoCE v2 or InfiniBand) bypasses the CPU, uses zero-copy GPU-to-GPU transfers via GPUDirect RDMA, and delivers far higher effective bandwidth with an order-of-magnitude lower latency.
Do I need InfiniBand, or is RoCE enough?
Both work. InfiniBand is a lossless, purpose-built fabric standard in DGX SuperPOD deployments. RoCE v2 runs RDMA over Ethernet and reaches comparable throughput when configured correctly with PFC and ECN, often the better fit for cloud, colo and bare-metal Kubernetes clusters.
Can you fix distributed training issues without changing hardware?
Often, yes. Many underperforming clusters are misconfigured rather than under-provisioned: NCCL falling back to TCP, PFC/ECN disabled, wrong IB_HCA selection, NUMA-misaligned processes. We frequently recover lost throughput using the existing NICs and switches.
How do you prove the improvement?
We baseline with nccl-tests and per-iteration timing before touching anything, apply the configuration changes, then re-run the same benchmarks so the difference in AllReduce time and scaling efficiency is measured, not asserted.
Struggling with multi-node training?

Let's look at your NCCL config and network fabric and tell you what's wrong.