Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Apache Iceberg: Multi-Engine Lakehouse, Spec to Production
Highest Rated
Hot & New
Rating: 4.9 out of 5(16 ratings)
174 students

Apache Iceberg: Multi-Engine Lakehouse, Spec to Production

Master the Iceberg spec, every catalog, all engines, and the discipline to run a real multi-engine lakehouse.
Last updated 7/2026
English
English

What you'll learn

  • Read the Iceberg spec — manifests, snapshots, metadata.json, commit protocol — as a working mental model.
  • Pick the right catalog (Polaris, Unity OSS, Nessie, Glue, HMS) using the Decision Spine matrix and defend the choice.
  • Ship a complete medallion lakehouse on Iceberg with Kafka CDC, Spark MERGE, and multi-engine readers.
  • Run production maintenance — compaction, expire-snapshots, orphan-cleanup, rewrite-manifests — with real cost numbers.
  • Implement WAP, branching, tagging, and time-travel rollback patterns auditors and SREs both sign off on.
  • Migrate Hive, Delta, or bare Parquet warehouses to Iceberg using the right strategy for the situation.
  • Diagnose the four most common Iceberg failures using a 10-step playbook.
  • Pass an 8-dimension architect's rubric: spec, catalog, writes, reads, maintenance, security, cost, troubleshooting.

Course content

24 sections116 lectures9h 44m total length
  • The Monday Morning Lakehouse Crisis4:12

    Follow Sarah’s crisis of three compute engines, Snowflake, Databricks, and S3, and learn how one storage, many engines, one truth architecture fixes cross-engine questions and aligns the data.

  • Why Hive Metastore Hit Its Limits3:28
  • ACID on Object Storage — The Hard Problem4:25
  • What Open Lakehouse Actually Means4:27
  • Decision Spine v1 — When Iceberg Fits3:51

Requirements

  • Comfortable with SQL and core data-engineering concepts (you've built or maintained pipelines/ETL).
  • Familiarity with at least one query/processing engine (Spark, Trino, Snowflake, or similar) helps but isn't required.
  • Basic cloud-storage literacy (S3/ADLS/GCS object stores); no specific account or install is mandated to follow along.

Description

Apache Iceberg won the table-format war. Now you have to actually run one in production.

Most tutorials show you a CREATE TABLE and call it a lakehouse. This course takes you all the way: the spec internals, every major catalog, all the write and read engines, and the production discipline that separates a weekend pilot from a platform 200 engineers depend on.

Across 23 modules and 116 lessons (~12.5 hours), you'll follow Sarah — a Head of Data Platform running shared Iceberg storage across three continents — as she makes the decisions you'll soon make yourself. You'll read the spec as a working mental model (manifests, snapshots, metadata.json, the commit protocol), then choose a catalog with a real Decision Spine matrix comparing Polaris, Unity OSS, Nessie, Glue, and Hive Metastore — and defend the choice.

From there you build: a complete medallion lakehouse with Kafka CDC into bronze, Spark MERGE plus write-audit-publish into silver, and multi-engine readers (Trino, Snowflake, DuckDB) on gold. You'll run the maintenance nobody teaches — compaction, expire-snapshots, orphan cleanup, rewrite-manifests — with real cost numbers, and master branching, tagging, and time-travel rollback patterns that auditors and SREs sign off on.

What makes this course different

  • Spec-to-production, not hello-world. You finish able to operate Iceberg under real load, not just create a table.
  • Decision-first. A six-version Decision Spine (format, catalog, storage cost, engine cost, migration, an 8-dimension capstone rubric) gives you a defensible framework, not vendor hype.
  • Genuinely multi-engine. Spark, Flink, Kafka Connect, Trino, Snowflake, DuckDB — one storage, many engines, the way real lakehouses are built.
  • War stories + a 10-step diagnosis playbook for the four failures every Iceberg team eventually hits.

You'll learn to

  • Read the Iceberg spec — manifests, snapshots, metadata.json, commit protocol — as a mental model.
  • Pick the right catalog with the Decision Spine matrix and defend it.
  • Ship a medallion lakehouse with Kafka CDC, Spark MERGE, and multi-engine reads.
  • Run production maintenance (compaction, expire-snapshots, orphan-cleanup, rewrite-manifests) with cost numbers.
  • Apply WAP, branching, tagging, and time-travel rollback patterns.
  • Migrate Hive, Delta, or bare Parquet to Iceberg with the right strategy.
  • Diagnose the four most common Iceberg failures with a 10-step playbook.

Who's teaching

Built by Snowbrix Academy and taught by Amit — a credentialed practitioner (SnowPro Core, 2x Databricks-certified) who builds production data platforms for a living. Every pattern here is one you can defend on Monday.

If you're a data engineer ready to own the lakehouse — to stop following table-format tutorials and start architecting and operating one — this is your course.

Who this course is for:

  • Data engineers and platform engineers building or operating a lakehouse on Apache Iceberg.
  • Senior engineers who can write SQL but want to own format, catalog, and engine decisions — and defend them.
  • Data architects and tech leads evaluating Iceberg vs Delta/Hudi or choosing a catalog (Polaris/Unity/Nessie/Glue/HMS).
  • Anyone migrating Hive, Delta, or raw Parquet warehouses to an open, multi-engine lakehouse.
  • Not for: absolute beginners with no SQL/data background, or those wanting a single-tool click-along tutorial.