Skip to main content

5 posts tagged with "rdma"

View All Tags

· 10 min read

The fabric looks non-blocking on paper. A k=8 fat-tree, full bisection bandwidth, sixteen equal-cost paths between any two pods, ECMP enabled everywhere to use them. Then the first real training job lands and AllReduce sustains something like 60% of line rate, with a handful of spine uplinks pinned and others barely warm.

The usual explanation is that all-to-all traffic floods the fabric with so many flows that some are bound to collide. That has the mechanism backwards. ECMP copes well with many flows. It falls apart precisely because a collective gives it almost nothing to hash on.

· 7 min read

If you've ever written a NicClusterPolicy manifest for the NVIDIA Network Operator, you know the pain: the same repository, version, and imagePullSecrets copied and pasted across every single sub-component. OFED driver, RDMA shared device plugin, SR-IOV device plugin, Multus, CNI plugins, IPAM plugin, NV-IPAM - each one needs its own repository: nvcr.io/nvidia/mellanox and version: network-operator-v25.7.0. Change the version during an upgrade, and you're editing 8+ places in the same YAML. Miss one, and you get a partially upgraded cluster with mismatched component versions.

We recently contributed a fix for this: global config support for NicClusterPolicy (PR #2070). It's now merged into the NVIDIA Network Operator, and this post explains the problem, the implementation, and why it matters for anyone operating RDMA-capable GPU clusters on Kubernetes.