
Each lesson is broken into sub-modules that cover all the concepts of Data Engineering. Under the resources menu of the modules, you can find the lab document links wherever applicable.
For practicing the labs, you would require Azure Subscriptions. You can reach out to us if you need an Azure Subscription or the SandBox environment for doing the demos by mailing to kishore@mycloudbeaver.com or via WhatsApp at +91 86672 45568.
Explore data engineering on Azure, including Data Lake Gen2, Synapse Analytics, and data pipelines that transform and move operational data into analytical warehousing for reporting with Power BI.
Explore Azure data lake Gen2 as storage with hierarchical namespace for big data analytics. Learn Hadoop compatibility and read/write with Spark, Databricks, Synapse, HDInsight, and lake database with scalable replication.
Learn to query and transform data in a data lake with Azure Synapse serverless SQL pools, using openrowset, bulk options, and external tables for CSV, JSON, and Parquet data.
Transform data in a serverless sql pool by creating external tables from selective data using select queries, then encapsulate transformations in stored procedures and orchestrate them in pipelines.
Transform files with a serverless SQL pool to query Delta and CSV data in a data lake Gen2, create external sources and formats, and run a stored procedure.
Create a lake database to query data lake gen2 files via a schema and tables that map to folder paths, using templates and a designer.
Transform data with Apache Spark in Azure Synapse Analytics by loading a CSV into a data frame, adding a year column, and writing parquet files partitioned by year.
Transform data with Spark in Synapse Analytics by reading CSV with inferred schema, splitting names into first and last, and saving partitioned Parquet data by year and month.
Explore how Delta Lake in Azure Synapse Analytics enables data versioning and time travel, supports streaming with Delta tables, and ensures ACID properties for reliable data engineering.
Analyze data in a relational data warehouse using Azure Synapse Analytics, compare star and snowflake schemas, and apply surrogate and alternate keys for cube-based insights.
Load raw data into a relational data warehouse using staging, dimensional, and fact tables; generate surrogate keys, manage type zero, one, and two changes, date dimension, and post-load optimization.
Execute end-to-end loading of csv data into a relational data warehouse using Azure Synapse Analytics, copying from data lake Gen2 into stage tables and building dimension tables.
Build a data pipeline in Azure Synapse Analytics to move, copy, and transform data using linked services, notebooks, Data Lake Gen2, and data flow transformations.
Design and run a data pipeline in azure synapse analytics to load a csv into a sql pool, transform with data flows, and perform lookups and upserts with lineage monitoring.
Explore running spark notebooks in an Azure Synapse pipeline by configuring notebook tasks with dynamic parameters, testing in the environment, and enabling DevOps automation for scheduled executions.
Learn to run an Apache Spark notebook inside a Synapse Analytics pipeline, create and import notebooks, and execute a pipeline that transforms CSV data to Parquet.
Enable azure synapse link for cosmos db to configure the analytical store, then connect containers via a linked service and query analytical data with spark or sql.
Connect Azure Cosmos DB with Synapse Analytics using Synapse Link, enable the analytical store, create an Adventureworks sales container, load data, and query via Spark and serverless SQL.
Get started with Azure Stream Analytics to ingest streaming data, apply windowed queries (tumbling, hopping, sliding, session, snapshot), and visualize in real time with Power BI and Synapse Analytics.
Ingest real-time streaming data from Event Hub through Azure Stream Analytics into Synapse Analytics using hot or cold path options, and load into Data Lake Gen2 for querying and visualization.
This lab streams real-time data from an application into Azure Stream Analytics and delivers it to a Synapse Analytics pool via an Event Hub, using a script to provision resources.
Visualize real-time data by integrating Azure Stream Analytics with Power BI to create streaming datasets, dashboards, and charts that update near real-time as new data arrives.
Explore how Microsoft Purview provides governance across on premises, cloud, and SaaS data with a data map, data catalog, and discovery through a central registry.
Integrate Microsoft Purview with Synapse Analytics to scan, catalog, and govern data across data lake Gen2 and Synapse assets, track lineage, and search pipelines.
Explore how Microsoft Purview governs data in Azure Synapse Analytics, connecting data lake Gen2, lake databases, and dedicated SQL pools, and accessing the data catalog.
Explore Azure Databricks and Apache Spark in notebooks for data engineering and machine learning. Learn about Databricks file system, Delta Lake, metastore, and SQL warehouse across tiers.
Create and manage Apache Spark clusters in Azure Databricks via the Databricks portal, run interactive notebooks, query data with Spark and PySpark, and visualize results with charts using matplotlib.
Create and use a Spark cluster in Azure Databricks, run notebooks, load CSV data with a defined schema, perform Spark transformations and SQL queries, visualize results with matplotlib and seaborn.
Learn to run Azure Databricks notebooks from Azure Data Factory by creating a linked service and using a notebook task to execute a notebook with dynamic parameters.
In this course, you will learn how to implement and manage data engineering workloads on Microsoft Azure, using Azure services such as Azure Synapse Analytics, Azure Data Lake Storage Gen2, Azure Stream Analytics, Azure Databricks, and others. The course focuses on common data engineering tasks such as orchestrating data transfer and transformation pipelines, working with data files in a data lake, creating and loading relational data warehouses, capturing and aggregating streams of real-time data, and tracking data assets and lineage.
You can become a data professional, a data architect, or a business intelligence professional by learning about data engineering and building analytical solutions using data platform technologies that exist on Microsoft Azure. This course will give you a flavor of end-to-end processing of big data in Azure.
As a candidate for this certification, you should have subject matter expertise in integrating, transforming, and consolidating data from various structured, unstructured, and streaming data systems into a suitable schema for building analytics solutions.
As an Azure data engineer, you help stakeholders understand the data through exploration, and build and maintain secure and compliant data processing pipelines by using different tools and techniques. You use various Azure data services and frameworks to store and produce cleansed and enhanced datasets for analysis.