
Course overview
Compare OLTP and OLAP databases to reveal how transaction-focused row-oriented systems differ from analytics-driven column-oriented data warehouses, shaping data modeling and query design.
Illustrates how a column store organizes data into blocks across nodes and slices, using block min and max metadata to prune scans and enable parallel query execution.
Explain how column stores place data: map rows to slices, compress columns into blocks with mean and max metadata. Show how queries scan slices and blocks.
Model data with thoughtful keys and sort keys to avoid full scans in a data warehouse like Redshift, enabling efficient, one-pass aggregations across months and days.
Learn to reduce full scans with sort keys, one pass aggregation, and selective use of compound or interleaved keys for hierarchical or independent filters.
Understand distribution key's role in enabling massive parallelism by hashing a distribution column to slices, ensuring even data spread and efficient joins, while avoiding filter or time-based keys.
Co-locate data with the distribution key to minimize data copy and enable parallel joins; use sorting to enable merge joins and restrictive filters.
Designing a data model for a data warehouse is fundamentally different from designing the schema for your primary database. A data warehouse is meant for supporting aggregation queries that touch a huge number of records and are very expensive. The scale in terms of number of queries is at least an order of magnitude less than the number of queries seen by a consumer-facing primary database. This distinction leads to a fundamental different data layout and indexing model in the database used in a data warehouse. Such databases are called OLAP databases.
This course will take you into the internal architecture of an OLAP database with specific focus on areas that help you how a query runs on an OLAP database. We will walk though the levers that an OLAP database provides to optimize your queries. We will then cover various scenarios that will help you design the data model for an OLAP database used in a data warehouse.
Throughout this course, we will use AWS Redshift as the OLAP database. However, the principles covered in this course are general and can be applied in any data warehouse including Snowflake or Hive. Some of these principles are also applicable in real-time column stores like HBase.