
Explore InfiniBand fundamentals powering modern data centers with high speed, low latency, and lossless communication, essential for NVIDIA certifications and AI infrastructure roles.
InfiniBand serves as the backbone of modern data centers by enabling a faster network, addressing data growth, distributed computing, AI and ML training communication, low-latency needs, and full compute utilization.
This lecture uses a road and bullet train analogy to explain why InfiniBand is built from the ground up for high-speed, low-latency data networks, unlike Ethernet's evolution.
Trace the evolution of InfiniBand from SDR at 10 Gbps to QDR, FDR, EDR, and HDR, culminating in NDR, XDR, and a forthcoming GDR around 1600 Gbps.
Connect CPU-based servers, CPU-GPU systems, and storage on an InfiniBand network with switches and Ethernet gateways as needed. Use a topology with multiple switches to provide redundant links.
Explore the software and hardware components of InfiniBand, including cables, host channel adapters, and switches, and understand how data plane and control plane, managed by the subnet manager, enable connectivity.
Explore InfiniBand physical connectivity by linking compute and GPU nodes from multiple vendors via host channel adapters or embedded adapters to an InfiniBand switch.
Explore InfiniBand cabling basics, comparing copper DAC and optical AOC cables. DAC is short-range, cheap, low-power but heavy and EMI-prone; AOC offers longer range, lighter weight, and EMI-free connectivity.
Explore two-node InfiniBand connectivity on Ubuntu, using a switch or a direct DAC cable, install and validate drivers with OFED or RDMA, and perform fabric discovery with a subnet manager.
Install the idverb and devinfo utilities to verify an InfiniBand device, inspect port states, GUIDs, and LID, and confirm subnet manager status before OpenSM setup.
Connect two systems with a cable, verify the mlx card and modules, and check ibv dev outputs; ports initialize only after a subnet manager configures the fabric (OpenSM or Nvidia).
Deploy a subnet manager to discover topology, assign lids, build routing tables, and optimize paths in an InfiniBand fabric. OpenSM serves labs, while UFM targets production.
Install OpenSM and set up a two-node InfiniBand fabric, start the subnet manager with systemctl, verify ports and lids, and install rdma-core utilities for diagnostics.
Install infiniband-diagnose to run ibstate and ibstatistics, then inspect two HCAs, note PCI location, hardware MT4099, port statuses, 40 gbps link, and base LID while subnet manager LID remains 5.
Explain GUID, a 64-bit hardware identifier, and how system image GUID groups systems while each HCA's node and port GUIDs, configured by the manufacturer, govern traffic and LED assignment.
Understand local identifier (lid), a 16-bit InfiniBand port address used for subnet-wide, fast routing in a flat network, distinct from GUIDs and assigned by the subnet manager.
Explore InfiniBand connectivity with IDPing, a client-server tool that tests node-to-node latency, routing, and subnet manager status for MPI and RDMA workflows.
Discover and map the InfiniBand fabric using IBNet Discover, identify switches, ports, and connected nodes, and visualize topology from a node to U1/U2 servers.
Explore InfiniBand utilities such as IB Link Info and IB TraceRT to troubleshoot networks, visualize topology, and trace paths between nodes, switches, and ports across the fabric.
Compare Ethernet and InfiniBand, contrast tcp/ip with rdma, and highlight trade-offs in speed, latency, cost, and ecosystem for high performance computing and AI workloads.
Compare tcp/ip and infiniband architectures, showing tcp/ip as layered and general purpose, while infiniband targets high-performance computing with rc and ud and zero-copy rdma.
Explore the NVIDIA InfiniBand stack, from the hardware and software that power robust data center networks to NVIDIA’s integration of InfiniBand with GPUs, RDMA, and AI fabric.
Explore NVIDIA's InfiniBand hardware stack from cables to Metro X long haul, detailing Linkex cables, ConnectX NICs with Vpi, DPUs like Bluefield, and quantum switches and gateways for scalable HPC.
Explore the NVIDIA InfiniBand software stack, including Mellanox OFED and DOCA OFED, UFM for telemetry and management, and NCCL and RCCL for GPU multi-GPU communication across InfiniBand and NVLink.
Explore OFED, the open fabrics enterprise distribution, and DOCA OFED for DPUs and bluefield devices. Verify drivers with lsmod and OFED info, noting kernel versus external modules.
Download the appropriate mlx ofed package from nvidia, extract it, and run mnx ofed install with add kernel support to install drivers and update firmware, then verify with ofed info-s.
Want to understand how modern AI really works behind the scenes?
This beginner-friendly course introduces you to InfiniBand, the high-performance networking technology powering AI data centers, GPU clusters, and HPC environments. If you're preparing for NVIDIA certifications or aiming for roles in AI infrastructure, cloud computing, or solutions architecture, this is a must-have foundational skill.
InfiniBand is designed for low latency, high throughput, and efficient GPU communication, making it essential for distributed AI training and large-scale machine learning workloads.
This course breaks down complex concepts using real-world analogies, visual diagrams, whiteboarding sessions, demos, and comparison tables — so you can learn faster and retain more.
What you’ll learn:
InfiniBand fundamentals: architecture, hardware, and software
Key components: HCAs, switches, subnet manager, and fabric design
InfiniBand connectivity and communication flow
Real-world AI and HPC use cases
InfiniBand vs Ethernet: performance and design differences
InfiniBand vs TCP/IP: protocol and communication model comparison
Whether you're a beginner, cloud engineer, solutions architect, or AI enthusiast, this course will help you build a strong understanding of high-performance networking for AI.
By the end, you’ll confidently understand how AI systems scale — and how InfiniBand enables that performance.
Start your journey from AI user to AI infrastructure expert today.