
Explore how generative AI fits into the Databricks Lakehouse, with Mosaic AI model serving, vector search, prompt engineering, and building AI agents for certification prep.
Databricks is a unified data intelligence platform that consolidates data engineering, analytics, and generative AI workloads, enabling ETL, data warehousing, and BI with Delta Lake, Unity Catalog, and Spark.
Explore the data estate evolution from data warehouse to data lake to data lakehouse, guided by ACID, Delta Lake, and enterprise-ready governance on the Databricks platform.
Understand how Apache Spark evolved from the Hadoop ecosystem to the Spark ecosystem, adopting an in-memory processing model with resilient distributed data frames and lineage.
Explore how Apache Spark distributes work across driver and worker nodes, interfaces with cloud data sources, and uses DataFrames, Spark SQL, and the machine learning library for distributed analytics.
Learn to work with Spark data frames, PySpark and SQL in Databricks lab 1, using Unity Catalog volumes and Azure Data Lake to manipulate CSV data.
Create and manipulate Spark data frames from csv files in Unity Catalog, define schemas, clean data, calculate tax, split customer names, and aggregate yearly sales by item.
Create delta tables from a csv in a unity catalog volume, enable versioning and lineage, and use the resulting delta table for AI, ML, and BI workloads.
Trace the medallion architecture from raw delta data through bronze, silver, gold layers using spark declarative pipelines to produce cleansed data and business aggregates for BI and AI workloads.
Implement the medallion architecture with spark by building bronze, silver, and gold delta tables in Unity Catalog, enforcing schema, and deriving year-wise sales aggregates.
Trace the evolution from artificial intelligence to generative ai and large language models, and explain how transformer-based foundational models power contemporary ai applications.
Learn how AI agents and compound AI systems combine a foundational language model with APIs and external tools to autonomously plan and execute tasks, delivering ROI beyond chatbots.
Explore prompt engineering by crafting precise prompts to guide AI agents and generative models. Learn elements: goal, context, expectations, and source for clear outputs across language, code, and image tasks.
Explore chain-of-thought prompting, zero-shot, and few-shot techniques, plus best practices for structuring prompts, breaking tasks down, and controlling LLM output with clear syntax and delimiters.
Learn to call a Databricks large language model with the OpenAI SDK, including setup, token configuration, and chat completions; explore multimodal image input with base64 encoding.
Perform a hands-on lab using the ai_query sql function to add sentiment and one-sentence summary to restaurant reviews. Build bronze to gold delta tables with medallion architecture in Unity Catalog.
Explore MLflow and PyFunc in Databricks to build, package, and deploy custom AI models using the Mosaic AI model serving framework and model registry.
Explore the Mosaic AI agent framework in Databricks, a unified system to deploy, govern, and evaluate models on enterprise data with Unity Catalog and endpoints.
Design and deploy a custom chatbot within Databricks using mlflow pyfunc, register it in unity catalog, and serve via mosaic ai, featuring a batman persona and a custom system prompt.
Publish version 1 of a custom chatbot to a Mosaic AI real-time endpoint via Unity Catalog. Configure compute type, traffic routing, concurrency, and usage logging to a catalog table.
Learn how retrieval-augmented generation grounds enterprise queries in private data using vector embeddings and a vector index within a databricks pipeline, enabled by mosaic ai, mlflow, and unity catalog integration.
Fine-tune a pre-trained large language model using task-specific prompt–response data to tailor outputs for a business use case, boosting accuracy, and explore retrieval augmented generation for dynamic enterprise knowledge.
Learn data preparation for retrieval augmented generation by chunking documents into a vector search index using fixed-size with overlap, semantic, recursive, adaptive, and context-enriched strategies.
Learn to call vector embedding engines in Databricks using the OpenAI SDK, exploring models in the Unity catalog and returning 1024-dimension embeddings.
Prepare data for a rag pipeline by extracting pdf and image content, building bronze delta tables in Unity Catalog, and finalizing a gold delta table for vector embeddings.
Compare three vector index types in Databricks: delta sync with Databricks-hosted embeddings, delta sync with self-managed embeddings, and direct vector access, highlighting CDC, pre-calculated embeddings, and rack pipelines.
Explore how a Mosaic AI vector search index using HNSW organizes vectors in high-dimensional space and enables semantic, keyword, and hybrid searches for retrieved documents.
Create a custom rag model by wiring a vector search index with Delta Sync in Unity Catalog and deploy the RAG pipeline using MLflow, Python, and Mosaic AI model serving.
Deploy a rag model behind a mosaic ai endpoint and enable real-time querying. Configure logging and an inferencing table in the unity catalog to track usage and request details.
Evaluate a rag model’s production performance with MLflow’s evaluators for safety, relevance, and groundedness; log request-response pairs to a delta table dataset and run post-deployment evaluations.
Discover how Langtian, an open source framework, enables building scalable ai agents within the Databricks ecosystem, simplifying rag pipelines, multi-agent orchestration, and observability with Langsmith.
Explore getting started with LangChain on Databricks: install the Databricks LangChain SDKs, invoke a Databricks hosted model, and build a Spark dataframe agent to query restaurantreviews.csv.
Build a unity catalog-hosted sql function, register it in unity catalog, and attach it to a langchain-enabled databricks agent to query a delta table of electronics products.
Build a sequential multi-agent framework with LangChain in Databricks to generate a 12-week Azure AI engineer study plan, mapping topics to MS Learn modules via a three-agent chain.
This course is a complete, exam-aligned guide to the Databricks Certified Generative AI Engineer Associate certification, designed for professionals who want to build, deploy, and manage Generative AI applications on Databricks with confidence.
Generative AI on Databricks goes far beyond prompt writing. To succeed in real-world projects—and in the certification exam—you must understand how foundation models, embeddings, vector search, RAG pipelines, MLflow, and governance work together. This course focuses exactly on those skills.
You will learn how to design and implement Retrieval-Augmented Generation (RAG) systems, use Databricks Vector Search for semantic retrieval, manage embeddings effectively, and integrate LLMs into scalable data and analytics workflows. Every concept is explained with a clear mental model, followed by hands-on demonstrations using Databricks-native tools.
The course is structured to closely align with the official Databricks exam blueprint, helping you understand not just what to do, but why it works—an essential skill for both certification success and real-world engineering.
In this course, you will:
Understand Databricks’ Generative AI architecture and ecosystem
Build end-to-end RAG applications using embeddings and Vector Search
Apply prompt engineering techniques for reliable and grounded outputs
Track, evaluate, and manage GenAI models using MLflow
Follow best practices for governance, security, and cost awareness
Prepare confidently for the Databricks Certified GenAI Engineer Associate exam
Whether you are preparing for certification or looking to upskill in production-grade Generative AI on Databricks, this course provides a structured, practical, and exam-focused learning path.