
Explore AI fundamentals and generative AI, then map compute, networking, storage, data pipeline, AI platform, security, and observability across cloud, hybrid, on-prem, and edge.
Compare machine learning with deep learning, highlighting feature engineering versus automatic feature learning, and map techniques like regression, CNNs, and transformers to CPU or GPU hardware for diverse data.
Learn how generative AI operates—from gathering data and preprocessing to training large transformer models and deploying them for fast, low-latency inference, with RAG, GANs, governance, and guardrails.
Explore how AI infrastructure unlocks business outcomes in data centers by optimizing networks, automating provisioning, enabling self-healing, strengthening security, and powering analytics, collaboration, and IoT with Cisco solutions.
Explore ai/ml compute clusters as scalable, gpu-driven systems, balancing model types—from pre-trained to custom and fine-tuned—with lossless fabric and placement across on-prem, cloud, and edge to optimize parameters and performance.
Explore how Jupyter notebook serves as an interactive front door for data science and machine learning, connecting to GPU back-ends and data sources to run experiments and prototype models.
Compare traditional cpu-centric data centers with gpu-centric ai infrastructure, contrasting north-south cpu traffic with east-west gpu traffic, and outline blocks: network, compute, storage, virtualization, orchestration, monitoring for a scale-out design.
Explore how network and compute form a lossless AI fabric, linking GPUs via NVLink, with storage, virtualization, orchestration, and monitoring to maximize utilization.
Identify where AI workloads run—cloud, on-prem, hybrid, and edge—and weigh data gravity, latency, cost, and compliance. Emphasize portability via containers and open formats to avoid vendor lock-in.
Master data sovereignty, compliance, and governance to drive policy-led AI design. Align infrastructure placement with data residency rules like GDPR and HIPAA.
Cisco AI solutions present AI PODs, AI Canvas, and Hyperfabric AI as a validated, pre-integrated stack for AI-driven operations and faster time to value.
Map AI's four homes—cloud, hybrid, on-prem, and edge. Explore GPU, fabric, and storage demands, model choices, and right-sizing for language models tied to infrastructure.
Design an AI-ready network for GPU servers in a high-speed fabric, with unified storage, security and segmentation, and automation, following the five-step plan: plan, design, implement, validate, optimize.
Translate workload requirements into buildable, measurable data-center designs and validate them against goals to deliver scalable, efficient AI infrastructure.
Show how AI workloads require high bandwidth, low latency, scalable and redundant architectures with a lossless spine-leaf fabric, rich visibility, and congestion management to sustain GPU training at scale.
Assess latency consistency and jitter versus bandwidth, differentiate redundancy from resiliency, and explain how lossless fabrics protect against packet loss that cripples RDMA and AI performance.
Compare copper DAC and optical fiber for short in-rack hops and longer AI workloads, and summarize choosing Ethernet with RoCEv2 or InfiniBand based on latency, cost, and ecosystem.
Compare copper DAC and optical cables, learn structured cabling for airflow and reliability, and choose between Ethernet with RoCEv2 and InfiniBand for AI data center networks.
Explore layer 2 and layer 3 in a modern AI fabric, using VXLAN over an L3 underlay with BGP and ECMP to scale, plus fog computing for edge intelligence.
Learn a practical five-step migration to a dedicated AI fabric alongside your current network. Validate lossless performance and use a phased rollout to safely scale AI workloads.
Apply a practical five-step migration to brownfield data centers: assess ai requirements, identify gaps, build a parallel ai fabric, validate lossless behavior, and phase deployment for scalable, lossless ai workloads.
Design scalable, resilient artificial intelligence networks by applying bandwidth and low latency requirements, transport options like ethernet with rocev2 and InfiniBand, spine-leaf underlay, vxlan overlay, and lossless fog computing.
Explore rdma foundations, detailing remote direct memory access, cpu-bypassing lossless ethernet fabric via rocev2, and one-sided read/write for ultra-low latency gpu-to-gpu AI workloads.
Explore Cisco Nexus 9000 intelligent buffering with adaptive, fair, precise, resilient ECN and PFC management for AI traffic, using AFD, ETRAP, and DPP to balance elephants and mice.
Cisco Nexus 9000 uses adaptive, fair, precise, and resilient buffering with per-flow ML insights to manage AI/ML traffic through ECN and PFC.
Discover how data center bridging exchange (DCBX) auto-negotiates end-to-end PFC and ETS settings across switches and NICs to ensure lossless RoCEv2 with DCQCN, ECN, ETRAP, and AFD.
Map the data pipeline from raw sources through ingest, process, store, and serve to AI training, while highlighting data prep as the bottleneck and the need for high‑performance, low‑latency infrastructure.
Explore how equal-cost multi-path routing (ECMP) uses 5-tuple hashing to evenly distribute traffic across spine-leaf networks, while advanced load distribution fixes elephant AI flows and prevents hot links.
Trace the data pipeline from raw sources through ingest, process, store, and serve, and see how data prep drives model quality. Learn phase-specific infrastructure demands that enable AI training.
Link AI/ML workloads to the network fabric using Nexus Dashboard Insights to monitor flow throughput, latency, and congestion, tying network behavior to training performance and GPU utilization.
Explore how RDMA-driven data transfer achieves near-zero overhead with RoCEv2 over routable Ethernet, and how a lossless stack—QoS, ETS, PFC, ECN, DCQCN, DCBX—ensures fair, low-latency AI networking.
Explore AI hardware from CPUs, GPUs, DPUs, and NVLink to enable AI workloads. Provision Cisco UCS C-Series and X-Series servers with Intersight policies for virtualization and storage.
Design a balanced ai server by mapping workloads to c-series or x-series, sharing GPUs with MIG, and using lossless fabric and Intersight to optimize utilization and TCO.
Assess hardware readiness for data center AI by understanding DPU offload benefits, GPU memory sizing, and NVLink versus PCIe connections, then meet the Cisco servers.
Map workloads to UCS C-Series or X-Series, share GPUs with MIG, and cluster nodes to scale training, underpinned by a lossless fabric and Intersight management to optimize utilization and TCO.
Intersight uses a policy-driven model to configure UCS, turning policies into reusable profiles deployed across servers. It covers domain configuration, power, storage, LAN, QoS, and NTP for end-to-end AI provisioning.
Define policies and bundle them into profiles in Intersight; map rdma and storage to lossless qos classes end-to-end from server to switch; enable synchronized ntp for troubleshooting.
Explore Cisco AI PODs, Hyperfabric AI, and AI Canvas for a validated, pre-integrated AI infrastructure. Understand hyperconverged infrastructure that scales by identical nodes, with FlashStack, GPT-in-a-Box, and Run:ai on UCS.
Learn how virtualization and containers optimize hardware use, with VMs and hypervisors alongside Docker and Kubernetes. See software-defined storage and networking that support scalable AI workloads.
Implement a tiered storage strategy for AI, balancing petabyte-scale capacity and high throughput with block, file, and object storage, Fibre Channel, FCoE, NVMe, NVMe-oF, and software-defined storage.
Deploy the AI-enabled fabric with Cisco orchestration tools, stand up a real AI service using retrieval-augmented generation, and monitor with telemetry and Splunk for proactive operations.
Explore Cisco orchestration tools across CPUs, GPUs, DPUs, NVLink, and Intersight managed UCS, QoS, and AI solutions to optimize storage, NVMe and NVMe over Fabrics, and software-defined storage.
NDFC applies an AI-optimized template for a uniform lossless fabric with RoCEv2, PFC, ECN, QoS, and buffer management, while contrasting ACI with APIC and the define once, apply everywhere approach.
Learn how Cisco's orchestration tools automate AI-ready fabrics with NDFC, APIC, Nexus Dashboard, Intersight, and Hyperfabric. See how define-once, deploy-everywhere enables faster, scalable, policy-driven compute and fabric management.
Explore the orchestration landscape: NDFC automates the NX-OS spine-leaf fabric and APIC controls ACI, while Intersight and Hyperfabric manage cloud-based UCS compute and the AI fabric for deployment.
Explore cloud-managed data center orchestration with Intersight and Hyperfabric AI for on-premises UCS hardware, applying the policy-and-profile model to automate fabric health and monitor unified operations with Nexus Dashboard Insights.
Deploy open-source gpt-style models locally with retrieval-augmented generation to keep data private and under your control, using kubernetes, containers, and GPUs for on-prem inference behind an api endpoint.
Learn why AI infrastructure monitoring is its own discipline and how to observe every layer with Cisco tools, turning telemetry into action through lifecycle management.
Learn to deploy open-source GPT models locally for privacy, security, cost efficiency, and low latency, and use the three-step RAG pipeline—retrieve, augment, generate—to ground answers, reduce hallucinations, and monitor operations.
Explore why AI monitoring is essential, what to watch across every layer, and how Cisco tools transform telemetry into action, with lifecycle management ensuring uptime and optimal AI outcomes.
Implement a continuous lifecycle to keep firmware and software current, reducing vulnerabilities and downtime. Follow validated, staged rollouts—plan, validate, stage, deploy, monitor—to manage AI disruption and ensure reliable benchmarking.
Benchmarking measures real performance under load to establish a baseline and validate design, detecting regressions across network, compute, storage, and ensuring apples-to-apples results.
Explore how synchronized time enables log correlation, why NTP and PTP matter, compare Splunk Enterprise and Splunk Cloud, and outline the SPL query pipeline of data set, commands, and output.
Orchestrate fabric, compute, and AI services via seven panels using NDFC, APIC, Intersight, and Hyperfabric; deploy on-prem GPT with a RAG pipeline and monitor with Nexus Dashboard, Intersight, and Splunk.
This course contains the use of artificial intelligence.
Artificial intelligence has moved from the lab into the data center, and the network underneath it is no longer ordinary. Training and inference workloads demand lossless fabrics, GPU-dense compute, high-throughput storage, and telemetry that actually tells you what is happening. This course teaches you how to design, build, and operate that infrastructure on Cisco — and prepares you to pass the 300-640 DCAI (Implementing Cisco Data Center AI Infrastructure) exam, a concentration toward your CCNP Data Center certification.
We start with the fundamentals: what AI, machine learning, deep learning, and generative AI actually are, the workloads they create, and where they run — cloud, hybrid, on-prem, and edge. From there we move into network design for AI: bandwidth, latency, scalability, and the non-blocking, lossless fabric requirements that make or break a cluster. You will go deep on the technologies that matter for the exam and the job — RDMA and RoCEv2, PFC, ECN, and ETS congestion management, intelligent buffers, and QoS on the Nexus 9000 family.
Then we cover compute and storage: CPUs, GPUs, DPUs, SmartNICs and NVIDIA BlueField, Cisco UCS C- and X-Series, UCS management through Intersight, and modern storage with NVMe and NVMe-oF. Finally, you will deploy and operate a fabric with Nexus Dashboard Fabric Controller, stand up open-source GPT for RAG, and troubleshoot with telemetry and Splunk.
And you will not just watch — you will practice. The course includes seven downloadable, step-by-step hands-on lab guides that walk you through the real workflows: standing up an AI cluster fabric with NDFC, measuring RoCEv2 workload performance, deploying open-source GPT for RAG, troubleshooting an AI/ML fabric with Splunk, and more — each ready to follow in your own lab or a free Cisco dCloud environment. You also get two full practice tests with 130 exam-style questions, every answer backed by a detailed explanation, so you can find your weak spots and walk into test day knowing exactly what to expect.
Every lesson is tight, visual, and exam-aligned. By the end you will be able to read an AI infrastructure requirement and design the Cisco solution that meets it — with confidence on test day and on the job.