Skip to main content

3 posts tagged with "infiniband"

View All Tags

· 10 min read

The fabric looks non-blocking on paper. A k=8 fat-tree, full bisection bandwidth, sixteen equal-cost paths between any two pods, ECMP enabled everywhere to use them. Then the first real training job lands and AllReduce sustains something like 60% of line rate, with a handful of spine uplinks pinned and others barely warm.

The usual explanation is that all-to-all traffic floods the fabric with so many flows that some are bound to collide. That has the mechanism backwards. ECMP copes well with many flows. It falls apart precisely because a collective gives it almost nothing to hash on.

· 8 min read

Every GPU cluster design review reaches the same slide: the nodes and the leaf switch are in one rack, nothing is further apart than about three metres, and somebody has to decide what goes in the cable order. The textbook answer is passive Direct Attach Copper, and the textbook answer is usually right - lowest power, lowest cost, one fewer active component per link.

The part that bites is the number people carry around in their head. "DAC does 5 metres" was true at EDR. At NDR it is not, and a rack of 800G links planned on that assumption produces a cable order that does not physically reach.