About This Course
<div>InfiniBand Deep Dive: Networking for AI Data Centres</div><div><br></div><div>Welcome! I'm here to help you truly understand InfiniBand — the high-performance fabric powering the world's most demanding AI and HPC environments.</div><div><br></div><div>As AI workloads explode in scale, the network is no longer an afterthought — it is the bottleneck. Slow fabrics mean idle GPUs, longer training times, and wasted investment in expensive compute. This course gives you the deep, practical knowledge to understand, deploy, and troubleshoot the technology at the heart of modern AI data centres.</div><div><br></div><div>What you'll learn:</div><div><ul><li>Why traditional Ethernet and TCP/IP fall short for AI workloads — and how InfiniBand solves latency, throughput, and CPU bottleneck challenges</li><li><span style="font-size: 1rem;">The full InfiniBand architecture — Physical, Link, Network, Transport, and Upper layers — with real-world analogies that make concepts stick</span></li><li><span style="font-size: 1rem;">RDMA, Zero-Copy transfers, Queue Pairs, Memory Registration, and GPUDirect RDMA — the core technologies behind high-speed GPU communication</span></li><li><span style="font-size: 1rem;">How the Subnet Manager works — LID assignment, topology discovery, routing table programming, and failover with Standby SM</span></li><li><span style="font-size: 1rem;">Traffic isolation using Partition Keys (PKey) — configuring Full and Limited membership across multi-tenant AI clusters</span></li><li><span style="font-size: 1rem;">Quality of Service (QoS) — assigning Service Levels (SL), mapping to Virtual Lanes (VL), and configuring bandwidth weights in OpenSM</span></li><li><span style="font-size: 1rem;">Routing algorithms in depth — MINHOP, UPDN, Fat-Tree, Adaptive Routing — and why Adaptive Routing is critical for elephant flows in AI workloads</span></li><li><span style="font-size: 1rem;">Congestion control, Credit-Based Flow Control, credit loops, and how to prevent fabric deadlocks</span></li><li><span style="font-size: 1rem;">Fabric monitoring and management at scale using NVIDIA Unified Fabric Manager (UFM), including Cyber-AI and RBAC</span></li><li><span style="font-size: 1rem;">Hands-on troubleshooting using ibdiagnet, ibtracert, mlxlink, ibstat, smpquery, and more</span></li></ul></div><div><span style="font-size: 1rem;">This course is packed with visual analogies, architecture diagrams, and practical troubleshooting scenarios that make even the most complex concepts click. Whether you're a network engineer, a cloud infrastructure specialist, or an AI platform team member, this course will give you the edge to design and operate high-performance AI fabrics with confidence.</span></div><div><br></div><div>No InfiniBand experience required — just bring your curiosity and your ambition.</div><div><br></div><div><span style="font-size: 1rem;">Let's get started!</span></div><div><br></div><div><span style="font-size: 1rem;">NOTE - For learners preparing for NCP-AIN Exam</span></div><div><br></div><div>This course covers 40–50% of the NCP-AIN exam domains, with a focused deep dive on InfiniBand. It is a valuable study companion to cover a significant portion of topics for the certification.</div>
What you'll learn:
- Describe the full InfiniBand architecture — from physical layer to upper-layer protocols like RDMA and GPU Direct
- Design leaf-spine InfiniBand topologies suited for large-scale AI and HPC cluster deployments
- Implement QoS using Service Levels (SL) and Virtual Lanes (VL) to prioritise AI and HPC workloads
- Monitor and manage InfiniBand fabrics at scale using NVIDIA Unified Fabric Manager (UFM)
- Diagnose and resolve common InfiniBand fabric issues using tools like ibdiagnet, ibtracert, mlxlink, and ibstat