AI Infrastructure
Built for the networks behind AI.
GPU clusters are only as fast as the fabric between them. Netlume understands GPU cluster networking — rail-optimized topologies, RoCEv2 and InfiniBand, ECN, PFC and DCQCN — and connects fabric behaviour to the workloads that depend on it.
Congestion signalled to RoCEv2 senders (DCQCN)
Illustrative product UI with simulated demo data. Not real customer data or statistics.
Why it is different
AI fabrics fail in ways traditional tools miss.
Most monitoring was designed for best-effort networks with averaged counters. AI fabrics need lossless behaviour, microsecond visibility and an understanding of the jobs on top.
Lossless by design
RDMA over Converged Ethernet expects a near-lossless fabric. PFC and ECN must be tuned together — get them wrong and you trade drops for pause storms.
Synchronized traffic
Collective operations make thousands of GPUs talk at once. Incast and microbursts appear in microseconds and vanish before a poll interval ends.
Tail latency is job latency
A training step waits for the slowest flow. One congested port or flapping NIC can slow an entire job without any alarm firing.
Scale and symmetry
Rail-optimized fabrics have thousands of identical high-speed links. Problems hide in small asymmetries — one ECMP member, one rail, one optic.
Congestion control
ECN, PFC and DCQCN — understood as one loop.
Lossless Ethernet relies on several mechanisms interacting correctly. Netlume models the whole loop, so it can tell whether a slowdown comes from load, thresholds, a misbehaving NIC or a physical fault.
Coverage
What Netlume understands in an AI fabric.
- GPU cluster networking
Back-end fabrics that connect GPU servers for collective communication, typically separate from front-end and storage networks.
Watches · Rail mapping, NIC-to-leaf adjacency, per-rail utilization, job placement.
- RoCE / RoCEv2
RDMA over Converged Ethernet. RoCEv2 runs over UDP/IP (port 4791), so it is routable across leaf-spine fabrics.
Watches · Lossless traffic class, DSCP/priority mapping, retransmits, out-of-sequence events.
- InfiniBand
A purpose-built, credit-based lossless interconnect widely used for HPC and AI clusters.
Watches · Port state, link errors, congestion counters and topology from fabric management.
- Ethernet AI fabrics
High-radix 400G/800G leaf-spine Ethernet built for AI, using RoCEv2, ECN and PFC with careful buffer tuning.
Watches · ECMP balance, buffer profiles, queue occupancy, link and optic health.
- ECN
Explicit Congestion Notification. Switches mark packets as queues grow instead of dropping them, signalling senders to slow down.
Watches · Marking rate per port and class against Kmin/Kmax thresholds.
- PFC
Priority Flow Control (802.1Qbb). Pauses a single priority on a link to prevent drops in the lossless class.
Watches · Pause frames and duration, pause propagation, storm and deadlock patterns.
- DCQCN
Data Center Quantized Congestion Notification. The rate-control loop combining ECN marks with CNPs returned to RoCEv2 senders.
Watches · CNP rates, sender rate reductions and recovery behaviour.
- Congestion & microbursts
Short-lived queue build-ups from synchronized flows, often invisible in averaged counters.
Watches · High-resolution queue telemetry, buffer watermarks, drop counters.
- Fabric topology
Leaf-spine and rail-optimized designs where GPU n of every server connects to rail leaf n.
Watches · Cabling against intended design, missing or asymmetric links, ECMP group membership.
- Latency
Collective completion time is bounded by the slowest path across the fabric.
Watches · Per-path latency, queueing delay, job step time correlation.
- Packet loss
Even small loss in the lossless class triggers retransmits that stall collectives.
Watches · Drops per class, FCS/CRC errors, optics, NIC retransmit counters.
- AI workload impact
The link between a network symptom and the job, tenant or model run it slows down.
Watches · Job-to-host mapping, step time, collective duration alongside fabric telemetry.
Investigation
From a slower training step to a specific queue.
When step time rises, Netlume works down from the job to the hosts, rails, leaves and queues involved — and shows which explanations it ruled out along the way.
Correlated telemetry · 20:15 – 20:45 UTC
Hypotheses · INC-4127
H1ECMP redistribution → TC3 congestion on leaf-07 Et12
92%supported · 7 evidenceH2Degraded optic or cabling on leaf-07 Et12
4%ruled out · 2 evidenceH3Host NIC firmware regression on GPU-WORKER-042
2%ruled out · 2 evidenceH4PFC storm / pause deadlock
2%ruled out · 3 evidence
Root cause candidate
evidence-linkedEast-west congestion following ECMP path redistribution resulted in queue saturation on leaf-07.
Confidence
Ruled out
- Optic / physical layer fault (Rx power nominal)
- Host NIC firmware regression
- PFC storm / deadlock
Illustrative product UI with simulated demo data. Not real customer data or statistics.
For AI infrastructure teams
Questions you can ask about your fabric.
Netlume is designed for heterogeneous AI infrastructure: Ethernet and InfiniBand fabrics, multi-vendor switching, SmartNICs and DPUs, and the Linux hosts at the edge of it all.
- ›Why did step time for run-0412 increase 23% at 20:31?
- ›Which rails are seeing PFC pauses above baseline right now?
- ›Is this all-reduce slowdown caused by the network or the hosts?
- ›Which GPU NICs show rising retransmits or link flaps this week?
- ›Did the last ECMP change unbalance traffic across spines?
- ›Are buffer profiles consistent on every leaf in AI Fabric A?
Make your AI fabric explainable.
Walk through a GPU-fabric investigation with our engineers and discuss how Netlume would connect to your environment.