
Learn how Nvidia technologies power AI infrastructure across deep learning, data science, machine learning, and generative AI, with hands-on guidance on GPUs, software, and Nvidia certification.
Explore NVIDIA certifications for infrastructure professionals and developers, focusing on the AI infrastructure and operation NCAA certification, including exam format, 50 questions, 1 hour, $125, and prerequisites.
Identify topics covered in the Nvidia certification, including accelerated computing, AI, machine learning, deep learning, GPU architecture, Nvidia software stack, exam study guide, and four-module weight distribution.
Explore the drivers of ai evolution, including data explosion, computational power growth with gpus and cloud resources, and algorithm breakthroughs like neural networks, transformers, and diffusion models.
Explore how industries apply ai technologies, from autonomous vehicles with real-time object detection and decision making to healthcare image analysis, video surveillance, finance fraud detection, and retail optimization.
Define ai as machine-simulated human decision making. Explain ml as learning without explicit programming, and dl as neural network based learning, then introduce generative ai that creates new content.
Explore a chess analogy to explain AI, ML, DL, and Gen AI: AI uses rules, ML learns from past games, DL learns by self-play, and Gen AI creates variants.
Learn how transformer models use attention to understand relationships between words, predict the next word, and scale content generation through parallel computing.
Explore the four core blocks of an ai centric data center: compute nodes, network, storage, and support infrastructure, and the power, cooling, and space constraints for high density gpu workloads.
Examine how data centers allocate power between IT equipment and cooling, lighting, and overhead, and learn the power usage effectiveness (PUE) metric—total energy over IT energy, targeting around 1.2.
Explore how ai centric data centers rely on compute power, from cpu to gpu, and how cuda enabled gpu compute unlocked breakthroughs like AlexNet for ai image recognition.
Compare cpus and gpus: cpus offer flexible, low-latency general purpose computing, while gpus provide high throughput with thousands of cores for parallel tasks like graphics rendering and AI.
Compare CPU and GPU architectures by detailing multi-core CPUs with L1, L2, and L3 caches and ALU and control units, versus GPUs with thousands of cores and dedicated GPU memory.
Trace Moore's law, once doubling transistors every 18–24 months, and show how lithography and rising costs slow momentum, while GPUs, AI accelerators, and chiplets boost performance.
Data processing units offload networking, storage, encryption, and security tasks from the central processing unit and graphics processing unit, enabling efficient data center and ai workloads with multi-tenant isolation.
Explore the four AI data center networks: compute, storage, in-band management, and out-of-band management, and why isolation improves performance, latency, security, robustness, and scalability.
Compare four network fabric types and their purpose, implementation, and design features, including InfiniBand, RoCE, and NVLink for compute networks, plus in-band and out-of-band management and storage fabrics.
Compare Ethernet and InfiniBand to understand high speed data center networking, highlighting Ethernet as general purpose and InfiniBand as a low latency high performance option.
Explore converged Ethernet (CE), a single Ethernet fabric that carries LAN, SAN, HPC, and RDMA over converged Ethernet with redundancy, high bandwidth (40–400 Gbps), low latency, and cost efficiency.
Explore how storage supports AI workloads in an Nvidia data center, detailing NVMe local storage, parallel file systems, NFS, and object storage, plus tiered lifecycle policies.
Explore cloud versus on prem GPU infrastructure, weighing low cost and scalable pay as you go resources against security, sovereignty, and full data control for compliant environments.
Trace NVIDIA’s evolution from gaming GPUs to AI and HPC, moving from GeForce 256 to cuda-enabled parallel compute and DGX systems.
Explore Nvidia's six-layer technology stack, from CPUs and networks to movement via NVLink, RDMA, and InfiniBand. Learn core libraries, gpu drivers, virtualization, monitoring, and Docker/Kubernetes integrations with TensorFlow and PyTorch.
Explore the physical layer building blocks and compare Nvidia RTX GPUs for gaming, workstation, and data center use, including real-time collaboration and virtual workstations.
Inspect an old Nvidia graphics card with gpu-z to identify shaders, cuda cores, tensor cores, and whether ray tracing is supported on a GTX 710.
Explore the DGX platform, a data center AI system with eight Nvidia GPUs, Nvswitch, and 15 terabytes of NVMe SSD, designed for training, inferences, and analytics at scale.
Explore how DGX SuperPOD scales AI workloads by interconnecting multiple DGX nodes into an exascale-class AI supercomputer. Learn about InfiniBand connections, storage, and multi-tenancy for enterprise, labs, and federated learning.
Explore ConnectX InfiniBand and converged Ethernet, powered by Nvidia and Mellanox, using host channel adapters (HCA) and DMA for ultra low-latency, high-throughput AI training and HPC interconnects across data centers.
Explore Nvidia BlueField super NICs, or DPUs, that offload networking, storage, and security from CPUs to dedicated chips, enabling zero-trust security, AI and HPC-optimized, multi-tenant data centers.
Explore NVIDIA reference architectures to kick-start fast, reliable data center designs using DGX and AI platform, with DGX A100 networking, white papers, and certified systems.
Explore how gpu cores, including cuda cores, tensor cores, and ray tracing cores, handle diverse tasks from versatile calculations to artificial intelligence and realistic rendering, with Nvidia drivers and api.
Compare CUDA cores, tensor cores, and RT cores, outlining their use cases from general computation and graphics rendering to AI training and inference workloads and real-time ray tracing.
Explore the DGX platform timeline, from DGX one (2016) through A100, Grace and Hopper, to Blackwell and DGX spark, highlighting CPU-GPU evolution and NVLink connectivity.
Explore Nvidia DGX deployment options from on-prem to pay-as-you-go cloud, and review DGX family, CPUs, GPUs, NVLink, PCIe Gen5, DGX OS, CUDA, cuDNN, NCCL, and SMI.
Compare DGX A100 and H100 systems, noting eight GPUs with A100 tensor cores and 80 GB per GPU (640 GB total), OS five and OS six, networking and management features.
Explore layer two data movement and i/o acceleration, covering nvlink, gpudirect, rdma, storage, and hpc fabric with infiniBand and opensim to enhance inter-device communication.
NVLink provides a high-speed direct connection between GPUs and CPUs, bypassing the PCI express bus for faster data transfers. Nvswitch enables scalable GPU-to-GPU communication in multi-GPU systems.
Explore InfiniBand, an open standard, a lossless, high bandwidth interconnect for HPC, with adapters, switches, and cables enabling end-to-end networking across distributed GPU clusters.
Compare Nvidia InfiniBand and converged Ethernet, where DGX SGX uses an sqs host channel adapter and Ethernet relies on NIC connections and spectrum switches. Latency, not speed, drives the choice.
Explore direct memory access and remote DMA concepts, showing how GPUs access system memory directly via DMA channels, bypassing the CPU over InfiniBand and RDMA over converged Ethernet.
Explore Nvidia's gpu direct dma and direct storage, enabling direct access to gpu memory across hosts via dma capable networks while bypassing the cpu and operating system.
Explore GPU direct storage, a direct channel from GPU to local NVMe storage that bypasses CPU and system memory, reducing IO bottlenecks and boosting training performance.
Compare GPUDirect RDMA and GPUDirect storage to see how low-latency GPU-to-GPU across nodes and high-bandwidth data loading from NVMe or parallel storage meet HPC and data workloads.
Discover layer three, where a DGX-optimized OS and GPU drivers enable virtualization for AI workloads, with DGX OS based on Ubuntu 22.04 LTS and includes CUDA toolkit and Docker.
Install and configure Nvidia gpu drivers to enable gpu utilization for HPC and graphics workload, and use Nvidia SMI to verify installation and read gpu information.
Discover gpu virtualization, where hypervisors create virtual GPUs and MiG across physical GPUs, enabling sharing across virtual machines with near-native performance in pass-through mode and flexible isolation.
Compare software-based vGPU and hardware-based MiG isolation, using a coworking analogy, and learn how hypervisors, Nvidia SMI, Linux, and containers enable an AIML workload.
Compare MiG and vGPU: hardware-level isolation with up to seven MiG per GPU and 5 GB slices, versus software-based virtualization for AIML and HPC workloads.
Explore layer four core libraries, focusing on CUDA as a compute architecture for parallel programming and GPU communication, enabling general purpose compute on GPUs with native C, C++, and Python.
Explore CUDA, Nvidia's GPU programming model, to bridge normal code with parallel GPU power and accelerate compute via frameworks like TensorFlow and PyTorch.
Install cuda easily by following Nvidia's docs, verify with Nvidia SMI, and run the toolkit on x86 or ARM GPUs across Linux or Windows with gcc prerequisites.
Explore how the NVIDIA collective communications library (NCCL) abstracts multi-GPU communication across NVLink, NVSwitch, PCIe, and RDMA, enabling efficient allreduce and broadcast with topology-aware optimization.
NVLink, NVSwitch, PCIe and RDMA form a hardware expressway for fast GPU transfers, while NCCL is a software library that organizes many transfers efficiently across GPUs.
Explore monitoring and management of a data center, using Nvidia SMI, data center GPU monitor, and base command manager to optimize resources, track temperatures, performance, and alerts.
Learn how Nvidia SMI monitors and manages a GPU on a single node, showing utilization, memory, power, temperature, processes, and clock speed via CLI, installed with the GPU driver.
Explore DCGM, the data center GPU manager for enterprise monitoring across multi-node clusters, tracking utilization, power, temperature, memory bandwidth, and error rates with the DXM exporter for Prometheus and Grafana.
Launch and govern your AI data center with base command manager, monitoring and orchestrating GPU resources and workloads through its cli, web ui, and rest api.
Use Nvidia SMI for single-system status, DCM for multi-node monitoring and alerts, Kubernetes GPU operator for cluster management, and OpenSM with the base command manager for enterprise AI data centers.
Explore Nvidia Clara for healthcare AI, Merlin for scalable, personalized recommendations, and Nims as an inference microservice to deploy AI models at scale across cloud or edge.
Explore the Nvidia solution stack from physical layers to GPU programming, including Docker, Kubernetes, CUDA, TensorFlow and PyTorch, with monitoring via Prometheus and Grafana.
Learn about Nvidia ai enterprise, the operating system for enterprise ai, and Nvidia ai factory that groups Nvidia services across cloud, data center, and edge.
Nvidia's AI factory turns data into intelligence by processing it with GPUs to train and deploy AI models. It supports on-prem, cloud, or hybrid deployments across the AI lifecycle.
Explore Nvidia's AI workflow: data processing, model training, optimization, and inferencing, with guidance on preparing data, training models, and deploying optimized solutions via Tensorrt and Triton Inference Server.
Explore how TensorFlow and PyTorch frameworks provide ready-made tools and building blocks for faster modeling, data handling, augmentation, cleaning, retraining, and model versioning with Nvidia optimization.
Nvidia differentiator lies in delivering optimized GPU libraries and cuDNN-backed acceleration, enabling PyTorch or TensorFlow to train and run models on GPUs fast while you focus on code.
This lecture contrasts model training and model inference, showing how training requires high compute, memory, and multi-GPU setups, while inference focuses on low latency, high throughput, and optimized models.
Compare training job scheduling with container orchestration for inferences, using air traffic control analogies. See how Slurm and Kubernetes manage resource allocation, scaling, and health checks for AI workloads.
Slurm handles resource allocation and batch AI training jobs, while Kubernetes manages container orchestration for always-on inferences and services, with GPU integration via CUDA-aware Slurm and Nvidia GPU operator.
Learn how Slurm coordinates multi-node training on DGX systems using DXM, NCCL, and GPUDirect RDMA, while Kubernetes uses the GPU operator and DQM exporter for Prometheus and Grafana metrics.
Understand MLOps through a kitchen analogy, then document, version data, automate deployment, and monitor models to ensure consistent, trustworthy predictions using Nvidia tools.
Learn why MLOps is essential for production deployment, automation, and monitoring, using CI/CD pipelines, event-driven retraining, and a model registry to manage versions.
Explore how Nvidia tools support every MLOps stage—from data preparation and feature engineering with RAPIDS and Nemo data curator to training, optimization with Tensorrt, and deployment with Triton.
Embark on a transformative journey into the world of AI infrastructure with this comprehensive course designed to prepare you for the Certified Associate: AI Infrastructure and Operations (NCA-AIIO) certification. Whether you're an IT professional, system administrator, or DevOps engineer, this course equips you with the foundational knowledge and practical skills needed to manage and optimize AI workloads in data center environments.
What You'll Learn:
AI Fundamentals: Understand the core concepts of Artificial Intelligence, Machine Learning, and Deep Learning, and their applications in modern computing.
GPU Hardware & Software: Gain proficiency in GPU architectures, including A100, H100, and B200, and explore essential software tools like CUDA, DCGM, and NGC Catalog.
Infrastructure Design: Learn about data center components, networking technologies such as NVLink and InfiniBand, and how to design scalable AI infrastructure.
AI Operations: Master the deployment, monitoring, and optimization of AI workloads in a enterprise data center, utilizing tools like DCGM, Slurm and Kubernetes.
Exam Preparation: Prepare thoroughly for the NCA-AIIO exam with detailed study guides, practice questions, and real-world scenarios. Gain a clear understanding of the exam objectives, learn tips to maximize your performance, and build confidence to pass the certification on your first attempt, validating your expertise in AI infrastructure operations