
Dremio is a modern data lakehouse platform that eliminates data movement and enables direct analytics on any source with an Apache Arrow powered sql engine and a virtual data layer.
Explore Dremio architecture with the Apache Arrow framework, Gandiva expression engine, Query Planner Intelligence, and reflection engine for high-performance, vectorized analytics and just-in-time code generation.
Compare Dremio and traditional data warehouses, highlighting direct analytics on data in place, a virtual data layer, and lakehouse advantages that reduce maintenance, lower costs, and speed time to insight.
Dremio acts as the unifying layer in the lakehouse ecosystem, bridging storage, compute, and analytics tools—from BI dashboards to data science and streaming—enabling unified data access and high-performance queries.
Install and run a local Dremio instance via Docker (or Kubernetes for production), access the UI on localhost:947, and enable data source configuration and query execution.
Create a dedicated AWS user for iceberg workflows with S3 full access and AWS Glue Console full access. Generate and securely store access keys for use in Python and PySpark.
Connect Dremio running in a docker container to AWS S3 as lakehouse storage by adding an S3 data source and configuring the warehouse path, selecting iceberg as the backend format.
Explore the Dremio query editor to create base tables in iceberg format, insert data, and run sql queries with schema evolution and time travel capabilities.
Master select queries in dremio, using Ansi standard sql to fetch specific columns or all columns from a product table, and apply upper or lower functions for case-insensitive data.
Discover how the distinct keyword in SQL removes duplicates and returns unique values from the category column, illustrated with a 25-row product table.
Explore essential aggregate functions in Dremio using product data, including count, min, max, and average price, and combine them in a single query.
Learn how to use aliases to rename columns and tables in sql queries, using as for name and id, create table aliases such as p, and handle alias errors.
Mastering data types and typecasting in SQL, this lesson guides you through creating an employees table, inserting data, and using cast to convert salary from string to float for calculations.
Master the concat function in Dremio to join two or more text values into a single output, with optional spaces, for readable product name, category, and price.
Learn how to use SQL upper and lower to standardize text and enable case-insensitive comparisons, with practical email examples and split_part to extract the domain.
Learn how the where clause filters data in SQL within a data lake, using equals, not equals, greater-than, and less-than conditions to retrieve only relevant stock records.
Explore how to filter string columns with sql queries using the where clause, including equals, not equals, and combining conditions with and or across product name and category.
Explore filtering boolean columns with the where clause in sql, using is active true or false, and combining conditions with and or to narrow results.
Explore filtering with between on a single column, using inclusive endpoints, with examples like price between 50 and 200 and order date between 1st January 2024 and 5th January 2024.
Explore how to filter data with the sql in clause, retrieving products by stock values (15 or 0) and by names (laptop, tablet, monitor) across numeric and text data types.
Learn pattern matching with SQL's LIKE operator to filter product names using wildcards like t% for starts with t, %s for ends with s, and _ for single characters.
learn to handle null and not null values in sql with dremio, including creating tables, inserting with nulls, and querying with is null and is not null.
Explore the group by clause in Dremio to analyze data by category and other columns, using aggregate functions such as max and sum to compute per-category totals.
Master the group by with having in SQL to filter aggregated categories, using the sum of prices and total stock, and learn why where cannot filter aggregates.
Learn to use a where clause with group by to filter product groups by price and is active, using selects on name, price, and is active.
Sort data with the order by clause in SQL, arranging product names alphabetically or prices in ascending or descending order.
Master the coalesce function to handle null values in sql queries, replacing missing emails with 'no email' and salaries with zero, improving readability and reliable calculations in data analysis.
Learn how to use the null if function in SQL to handle conditional null values, replacing salaries with null and excluding specific numbers from reports, with an employee table example.
Explore inner queries and subqueries to filter data dynamically, find specific products by product ID or name, and use results to match prices and sold product IDs.
Master the inner join in SQL by combining orders and product tables using on o.product_id = p.product_id, with aliasing, selecting specific columns, and applying where electronics to optimize queries.
Practice left join by combining the product and orders tables, returning all product records and nulls for unmatched orders, using aliases p and o with the on clause.
Explore the right join in sql by linking product and orders tables with aliases p and o and the on clause, showing product rows and nulls for non matches.
Explore how the union operator in SQL merges results from multiple tables, removes duplicates automatically, and consolidates data from employees and updated employees using single and multi-column queries.
Explore the union all operator in SQL to combine results from two tables, including duplicates. See practical examples with employee data and multiple columns like id and name.
Trace the evolution from traditional data warehouses to data lakes and data lake houses, comparing schema on write and schema on read, with governance and multi-engine analytics.
Trace the evolution from traditional data warehouses to data lakehouses, and learn how Iceberg delivers acid transactions, schema evolution, and unified metadata across processing engines in a lakehouse architecture.
Trace the shift from traditional databases to modern data lakes as iceberg delivers metadata management, metadata of metadata, acid transactions, and time travel across engines spark, flink, and duckdb.
Apache Iceberg is an open source table format built on Parquet or ORC that enables scalable analytics with a layered metadata system, supporting schema evolution, atomic operations, and time travel.
Explore Iceberg features like schema evolution, acid transactions, partitioning, and time travel, and see how they enable scalable, low-cost data analytics on large data lakes.
Explore how iceberg's metadata layer atop storage enables time travel, schema evolution, and acid transactions across spark, flink, duckdb, and trino.
Install pi iceberg, a Python library for Apache iceberg tables, in a clean virtual environment. Configure a local SQLite catalog and test with a Python script.
Install and start the Jupyter notebook server in the Pi Iceberg environment, then load a local Iceberg catalog and print its properties in an interactive notebook.
Install iceberg, polars, and duckdb from a Jupyter notebook using pip across Linux, macOS, and Windows to enable reading and writing iceberg tables with fast data frames and SQL analytics.
Discover how Iceberg catalogs manage metadata and coordinate operations across local and cloud deployments, from local folder to aws s3, Glue, Nessie, Minio, and Postgres catalogs.
Learn how to query Apache Iceberg tables using a local Iceberg catalog, select specific columns with scan and select, and convert results to pandas data frame.
Learn to create an Apache Iceberg table in Python using PyIceberg, connect to a local catalog, define a schema, insert data with PyArrow, and verify results with a table scan.
Filter Iceberg tables at the storage layer using a local catalog to load the sales product table, then apply row filter expressions like equal, greater than, in, and not in.
Connect iceberg with a SQL catalog that stores metadata in SQLite while the actual data resides in S3, enabling scalable cloud storage with local development.
Explore how Apache Iceberg operates in production with Spark, local machines catalog, and AWS S3, coordinating metadata and data layers across local and cloud storage for scalable analytics.
Install Pi Spark on your local machine from a Jupyter notebook to use the Python API for Apache Iceberg, enabling read, write, and schema management of Iceberg tables.
Configure a Spark session with Apache Iceberg, enable Iceberg extensions, and create a local Iceberg catalog and database to explore databases and a sample table in development.
Create iceberg tables in a local spark session, define a product table with six columns, insert data, and query it with spark sql to explore iceberg features like time travel.
Set up a Spark session with Iceberg integration, enable extensions, and configure a local Iceberg catalog with Hadoop storage and a Spark warehouse; explore databases and the local catalog.
Create an iceberg table with spark sql, insert 25 product rows into a local catalog, and leverage iceberg features such as schema evolution, partition evolution, and time travel.
Programmatically create iceberg tables from a spark dataframe by configuring iceberg, defining a schema, and saving as an iceberg table; read back to use schema evolution and time travel.
Read Iceberg tables with Spark's data frame api to query a local catalog and load data into a frame, then select product name, product id, and stock.
Set up a Spark session with Iceberg integration and load the product table. Apply aggregate functions like sum, max, min, and count, and use distinct to reveal inventory insights.
Configure a Spark session with Iceberg integration, load an Iceberg table, and apply PySpark filters using column expressions for equality, greater than, name matches, in lists, and negation.
Explore inserting data into Iceberg tables with Spark, using a data frame and the write API to append records. Learn copy-on-write behavior, snapshots, and optional time travel.
Learn how to perform update and delete operations on Apache Iceberg tables with Spark, using copy-on-write semantics, data lifecycle management, and time travel for auditing.
Discover time travel in Apache Iceberg by querying historical snapshots with as of version, enabling auditing, debugging, and restoring data from previous states.
Connect Dremio to a Postgres database via Docker, create a private network, start containers, add a Postgres data source named sales, and run queries to access 25 products.
Explore how to upload and organize csv data in Dremio, enable automatic schema detection, and query the Superstore dataset using the sql editor for fast data analysis.
Explore parquet files with Dremio, upload the Titanic parquet dataset, and experience columnar storage and optimized processing that enable fast analytics and a clear schema.
Compare physical data sets, which connect to external data without copying, with virtual data sets that use SQL logic to enable business-ready analytics in a four-layer lakehouse.
Learn how to create reusable views in Dremio to simplify filtering and share datasets. Build a consumer_segment view from the superstore data, enabling consistent analytics across teams.
Mastering Dremio: The Beginner's Guide
Unlock the power of modern data analytics with Dremio – your ultimate companion for high-speed SQL, data lake queries, and BI integrations!
What You’ll Learn:
Set up and configure Dremio on your system or in the cloud
Connect to diverse data sources like S3, SQL databases, and more
Create and manage datasets, reflections, and virtual datasets (VDS)
Write fast, scalable SQL queries on data lake storage
Integrate Dremio with BI tools like Superset
Understand user roles, permissions, and security
Learn best practices for performance tuning and optimization
Get hands-on with real-world examples and use cases
Why Learn Dremio?
Dremio is a game-changer in the world of data analytics and lakehouses. It lets you run lightning-fast SQL queries directly on your cloud data lake without moving data. Say goodbye to complex ETL pipelines and hello to self-service analytics!
Whether you're a data analyst, engineer, or BI professional – mastering Dremio will elevate your skillset and boost your productivity in the modern data stack.
Who This Course Is For:
Beginners with no prior experience in Dremio
Data engineers and analysts exploring data lake technologies
BI developers looking for fast and scalable query engines
Students and professionals entering the data analytics space
Anyone curious about the future of data lakehouses
Tools & Technologies Used:
Dremio Community Edition (local & cloud)
AWS S3 (for data lake examples)
SQL (basic to intermediate)
Apache Superset (for visualizations)
Sample datasets for hands-on labs
Course Features:
Step-by-step setup guides
Hands-on labs and projects
Real-world datasets
Quizzes and practice exercises
Downloadable resources
Lifetime access & updates
What Students Say About Our Courses:
"Clear explanations, practical examples, and great pace – just what I needed!"
"Perfect for beginners. Learned a lot about modern data stack tools."
"Another excellent course by the instructor. Highly recommended!"
Whether you're just getting started or want to future-proof your data skills, this course will take you from zero to confident Dremio user in no time.
Enroll now and start mastering Dremio today!