Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Databricks Data Engineering with AWS
Hot & New
New
Rating: 5.0 out of 5(9 ratings)
141 students

Databricks Data Engineering with AWS

Build a production lakehouse & deploy a real capstone project with Unity Catalog, Delta Lake, Lakeflow, DABs & CI/CD
Last updated 7/2026
English
English [Auto],

What you'll learn

  • Set up and govern a production Databricks workspace on AWS using Unity Catalog
  • Master Delta Lake internals — ACID transactions, time travel, constraints, and performance tuning
  • Design Medallion Architecture pipelines, first manually, then declaratively with Lakeflow Declarative Pipelines
  • Ingest data at scale with Lakeflow Connect — SaaS connectors, database CDC, and Auto Loader
  • Orchestrate production pipelines with Lakeflow Jobs — DAGs, retries, control flow, REST API and CLI
  • Build a complete production lakehouse for a real e-commerce business, from ingestion through five gold-layer outputs
  • Write unit and integration tests for Databricks pipeline code with pytest
  • Package and deploy pipelines using Databricks Asset Bundles (DABs)
  • Build a CI/CD pipeline with GitHub Actions that tests, validates, and deploys to a UAT environment

Course content

17 sections103 lectures25h 51m total length
  • About the Course5:15
  • Course Prerequisites3:58

    Identify required skills for this course, including Apache Spark basics, data frames, Spark transformations, Spark SQL, Python, and SQL, and note Databricks on AWS prerequisites and free-tier options.

  • How to access Course Material and Resources3:04

Requirements

  • Working knowledge of Apache Spark — DataFrames, transformations, and basic Spark SQL
  • Comfortable writing Python and SQL — both are used throughout the course and capstone
  • An AWS account (a free-tier account is enough to start; later chapters and the capstone incur modest AWS/Databricks usage costs)
  • No prior Databricks experience required — the course builds this from the ground up
  • Basic familiarity with Git and the command line helps in the CI/CD and DABs modules, though it isn't required going in

Description

Databricks has become the default lakehouse platform for data engineering on AWS — over 60% of the Fortune 500 run on it. But knowing individual features isn't the same as being able to design, build, test, and deploy a real production pipeline. This course takes you through both: you'll master the core Databricks and AWS skills chapter by chapter, then apply every one of them to a single, realistic capstone project — an end-to-end lakehouse built for a real business, deployed the way production teams actually deploy.

Learn Databricks on AWS and Build a Real Production Lakehouse from the Ground Up

  • Set up and govern a Databricks workspace on AWS with Unity Catalog from day one

  • Master Delta Lake — ACID transactions, time travel, constraints, performance

  • Design Medallion Architecture pipelines with Lakeflow Connect and Lakeflow Declarative Pipelines

  • Orchestrate production pipelines with Lakeflow Jobs — multi-task DAGs, retries, parameterization

  • Build a complete lakehouse for a real e-commerce business — five source systems, five business-critical gold outputs

  • Test, package, and deploy your pipelines with pytest, Databricks Asset Bundles, and GitHub Actions CI/CD

A complete path from Databricks fundamentals to a deployed, production-grade lakehouse — built one real skill at a time.

Phase 1 — Foundations. You'll start with the core skills every Databricks data engineer needs on AWS:

  • Workspace setup and Unity Catalog governance

  • Delta Lake internals — ACID transactions, time travel, constraints

  • Medallion Architecture, built by hand first, then declaratively with Lakeflow Declarative Pipelines

  • Ingestion with Lakeflow Connect — SaaS, database CDC, and Auto Loader

  • Orchestration with Lakeflow Jobs — DAGs, retries, control flow, REST API and CLI

Phase 2 — The Capstone. Every skill above gets applied to one continuous project: StepRight, a mid-size online footwear retailer with five source systems feeding five gold-layer outputs — daily revenue, customer 360, product performance, funnel analysis, and fulfillment health.

You'll ingest CDC and file-based data at production scale, then go further than most courses do:

  • Write unit and integration tests for your transformation logic

  • Package the project as a Databricks Asset Bundle

  • Wire up GitHub Actions CI/CD — test, validate, deploy to UAT

This is the same workflow real data platform teams run — not a toy example.

By the end of this course, you'll have built and deployed a governed, tested, production-structured lakehouse — end to end, on your own.

You'll walk away with:

  • A complete, working lakehouse project you built and can show, not just watched

  • Hands-on notebooks for every chapter, ready to import into your own Databricks workspace

  • A full GitHub repo structure from the capstone, showing exactly how a production project is organized

This isn't a features tour. It's the architecture, tooling, and deployment discipline real data platform teams run.

Disclaimer: This course was developed with the assistance of AI tools for content research, editing, and slide production. All technical content has been reviewed, tested and validated by the instructor.

Who this course is for:

  • Practising data engineers who already know Spark and Python and want to move from writing pipelines to running them in production
  • Data engineers and analytics engineers looking to add Databricks and AWS to their skill set with real, hands-on practice
  • Engineers who want to see an industry-standard, end-to-end lakehouse project — including testing, DABs, and CI/CD — built from scratch
  • Anyone already working with Databricks on AWS who wants to see it applied to a full production-style project