
Master Amazon Redshift with a three-part course covering history, architecture, and basics, plus hands-on demos, labs, and Tableau visualization tips for rapid adoption.
Rich morrow brings two decades in cloud and big data, guiding a hands-on Amazon Redshift course with AWS and Hadoop tooling.
Explore what a data warehouse is, its history, and its enterprise purpose. Compare OLTP and OLAP, and learn ETL and data marts for fast, scalable analytics.
Explore the high level architecture of a data warehouse, compare columnar versus row storage, and learn how star schemas with fact and dimension tables enable efficient analytics and compression.
Materialize data into multi-dimensional cubes for fast queries across customer, product, and time, while recognizing exponential growth and OLAP licensing limits as dimensions expand.
Explore the costs and maintenance overhead of self-owned data warehouses, including licensing, hardware, and vendor lock-in, and contrast with scalable public cloud options like Redshift.
Discover the benefits of a public cloud data warehouse, including lower costs, infinite scale, workload customization, and the managed service model with on-demand provisioning and elastic growth.
Explore Redshift benefits, including low pay-as-you-go costs, reservations that reduce long-term expenses, automatic compression, columnar storage, zone maps, and fast parallel loading.
Explore the benefits of Amazon Redshift, including easy setup with a web GUI and API access, automatic fault tolerance, and incremental backups to S3 for simple disaster recovery.
Explore Redshift security as an onion of layers, including IAM access control, VPC, security groups, encryption at rest, native permissions, auditing, and CloudTrail logging.
See how Pinterest and Nasdaq use Amazon Redshift to speed analytics, from data sources like Kafka, Hive, Hadoop, and a data lake on S3, delivering real-time KPI dashboards.
Nasdaq migrated from a legacy warehouse to redshift, cut costs 43%, secured data with vpc, direct connect, ssl, certificate signing, and at-rest encryption, while handling 14 billion rows per day.
Explore third-party visualization tools to turn diverse data sources into interactive dashboards with KPI insights and visualizations from Tablo, MicroStrategy, SAS, IBM, and Microsoft.
Compare Redshift with rdbms, commercial data warehouses, and Hadoop, including Presto, and clarify when Redshift is complementary or cannibalistic for olap vs oltp workloads.
Compare Amazon Redshift against RDBMS, cloud-based EDWs, in-memory engines, NoSQL engines, and columnar databases, and examine Presto's ability to query data where it lives.
Describe the architecture of Amazon Redshift, highlighting the leader node and compute nodes. Examine how data ingestion flows directly to compute nodes and how parallel execution drives performance.
Explore data loading options for Amazon Redshift, optimize parallel loads across slices, and apply the copy command from S3 while using data distribution, slicing, and vacuuming concepts.
Explore data distribution concepts in Redshift, including even, key, and all distribution types, to maximize parallel execution, minimize cross-node data transfer, and co-locate join data for efficient star schema replication.
Explore practical redshift usage from basic syntax and expressions to core copy, ddl/dml commands, data types and functions, and user and group permissions.
Explore Amazon Redshift basic usage, focusing on primary keys, distribution keys, and sort keys (interleaved and compound), plus copy and insert commands, and essential analytics and table management.
Create an AWS account to spin up a redshift cluster and sign into the console. Set up Virginia region CloudWatch billing alarms for five and ten dollars.
Create a multi-node Amazon Redshift cluster, configure the VPC security group for external access, create an IAM user and credentials, and learn to stop, restart, and snapshot to control costs.
configure sql workbench to connect to your redshift cluster, enable auto commit, set up the redshift driver, and run basic queries with a keep-alive connection script.
Create and configure an S3 bucket, upload sample data, and load into Redshift, inspecting the five tables—date, customer, parts, supplier, and the line order fact table—for a data loading lab.
Load data into a Redshift cluster by configuring a data loading script with credentials, creating tables, and preparing copy commands for part 2 using S3 and sequel work bench.
Complete part two of the lab by loading data into a Redshift cluster with copy commands, debugging outputs, and monitoring large data loads.
Continue loading data into a Redshift cluster from S3, monitor progress, debug errors, and validate data with selects and distribution checks; then snapshot the cluster.
Practice querying a Redshift cluster, view the queries library, validate data with counts, and optimize queries using create table as select and distribution keys.
Optimize a Redshift cluster by choosing sort and distribution keys for joins, load large tables, and compare query benchmarks to reveal performance gains.
Master data loading best practices for Redshift, including staging to S3, splitting data, compression, copy and vacuum tuning, and debugging load errors.
Master data loading in Redshift with copy from S3, vacuuming, post-load verify, and analyze statistics. Tune WLM slots, monitor data skew across slices, and handle copy errors.
Learn to tune Amazon Redshift query performance through thoughtful table design, choosing sort and distribution keys, managing compression, and maintaining leader node statistics with vacuum and analyze updates.
Tune Redshift query performance by choosing optimal distribution keys and join keys in a star schema. Use compression, the copy command, and explain to benchmark and optimize your queries.
Connect Tableau Desktop to your Redshift cluster, configure credentials, and build joined tables (public schema, line order, and customer key) to create interactive visualizations for dashboards.
Learn to explore Redshift data with Tableau, create region and revenue visuals, calculate margins, drill down by nation, and save or refresh worksheets.
Explore how Redshift fits into enterprise data warehouses, from self-owned to public cloud solutions, and master basic and advanced usage, including tuning and optimization.
In this Hands-on with Amazon Redshift training course, expert author Rich Morrow will teach you everything you need to know to be able to work with Redshift. This course is designed for the absolute beginner, meaning no previous knowledge of Amazon Redshift is required.
You will start with an introduction to the course, learning what a data warehouse is, the benefits of Redshift, and how Redshift compares to other analytics tools. From there, Rich will teach you the basics of Redshift, including data loading, data distribution concepts, and basic Redshift usage. Finally, this video tutorial will cover advanced topics, such as data loading best practices and tuning query performance.
Once you have completed this computer based training course, you will have learned everything you need to know to get started with Amazon Redshift.