
Explore data engineering technologies for beginners on Azure, understand course updates and topics, and learn how to access resources, lifetime access, and certification options.
Activate a sandbox from Microsoft Learn to keep using the Azure portal for free after the 12-month period, with sandbox resources deleted after four hours.
Encourage learners to provide reviews and feedback to support the course community, and learn to customize video quality, captions, and playback speed for an optimal viewing experience.
Explore Microsoft cloud services, understand portal and service categories, and provision a security service while managing subscriptions, resource groups, billing, and access through SSME and SQL Server Management Studio.
Create a free Azure subscription, get $200 credit for 30 days, 25 always-free services and 12-month offers, and sign up via a portal with phone verification.
Explore the Azure portal as the primary web interface to create, manage, and monitor resources with customizable dashboards, global search, and in-portal cloud shell.
Explore the full spectrum of Azure services, covering storage, databases, analytics, compute, security, and more. Learn how categorization, pricing models, and hybrid cloud options help tailor solutions for real-world applications.
Explore the difference between managed and unmanaged data services, focusing on ownership, patching, monitoring, backups, security, and regulatory compliance. Prefer managed services to reduce administration and overhead.
Explore azure's management group hierarchy to organize subscriptions and resource groups, apply policies and budgets, and manage costs across resources.
Create and manage azure resource groups by choosing a subscription and name, add resources from multiple regions, and understand that deleting the group removes all contained resources.
Explore your Azure subscription in the portal to view resource providers and their auto-registration. Register or unregister providers like Data Factory to access updated features.
Create and apply tags as metadata to Azure resources and resource groups to describe and organize them, then search by tag to view related resources and costs.
Azure identity and access management and role-based access control to delegate resource administration via scope-based role assignments, from management groups to individual resources, using roles like contributor, owner, and reader.
Demonstrates provisioning an Azure SQL database service from scratch, configuring a new server, resource group, and a sample database, with networking, elastic pools, and cost estimates.
Delete unused resources and resource groups to prevent costly bills, then set budgets, alerts, and cost analysis in Azure to monitor spending and forecast.
Gain familiarity with cloud computing by exploring the portal, creating services and dashboards, and managing subscriptions, resource groups, tags, and role-based access to build data solutions.
Discover the data engineer role, responsibilities, and the technologies they use, and why this profile is in demand, with a look at data science distinctions and modern data flow architecture.
Understand the data flow lifecycle from collection to analysis, including ingestion, cleaning, transforming, and modeling, and distinguish data engineer roles from data scientists and analysts.
Explore Azure data engineering technologies, from data lake to Data Factory v2 and Azure Databricks, that ingest, transform, and prepare data for analysis in Cosmos and data warehouses.
Clarify your role as a lead engineer, why this position exists and is in demand, and your responsibilities to render data for data centers using essential technologies.
Introduce Azure SQL database offerings, compare IaaS and PaaS options, and outline deployment options—single database, elastic pool, and managed instance—plus purchasing models, service tiers, and security features.
Explore Microsoft database offerings, including infrastructure as a service and platform as a service, and compare deployment options such as single database and elastic pool to understand strengths and trade-offs.
Choose this cloud relational database as a service for mission-critical workloads, offering a fully managed SQL Server engine with elastic pools, 99.99% uptime, zero replication, and seamless app integration.
Compare Azure IaaS virtual machine SQL Server and Azure PaaS database as a service, weighing responsibility for maintenance, patching, backups, and security. Explore availability, scaling, and automated management benefits.
Discover sql server paas deployment options, including single database, elastic pool, and managed instance, with isolated resources, shared pools, and dedicated sql server environments.
Explore deploying a sequel database server in a virtual machine and the lift-and-shift deployment option, with a demo to provision a database in a virtual machine.
Learn how SQL Server runs on Azure virtual machines, including automated patching, automatic backups to blob storage, and high-availability options with availability groups and templates.
Provision a SQL Server on an Azure virtual machine using a Microsoft image with a free developer license, then configure networking and connect with SQL Server Management Studio.
Explore the single database deployment option and its use cases in the Azure database service. Compare two purchasing models and service tiers—general-purpose, business critical, and hyperscale.
discover how to provision a single database in a logical server, compare deployment options: single database, elastic pool, and Mendy's instance, and configure firewall rules for access.
Compare data-based and vehicle-based purchasing models, with deployment options single database, elastic pool, and managed instance, and service tiers general-purpose, hyperscale, and premium for azure sql databases.
Contrast SQL database for online transaction processing with data warehouse for online analytical processing. Note horizontal partitioning across 60 compute nodes enables massively parallel processing and cost-saving cloud scaling.
Explore the elastic pool deployment option for a second database service, provisioning an elastic pool with multiple databases, configuring them together to save money and understand when it makes sense.
Azure elastic pool lets a collection of databases share compute resources, enabling bursts within a common pool. This approach reduces costs compared with single databases.
Demonstrates provisioning multiple Azure databases, creating an elastic pool, and moving databases into the pool, then scaling compute and storage with pricing options.
Explore managed instances on the Adjure sequel database service, covering use cases, migration, and create, update, delete operations, plus provisioning a second database management instance and a private network vm.
Learn about the Azure managed instance deployment option for SQL databases, enabling 100% compatibility with on premises SQL Server, secure virtual networks, and easy lift-and-shift migrations for large databases.
Compare on-premises SQL Server with Azure SQL Managed Instance, highlighting built-in high availability, authentication differences, automatic filegroups and in-memory OLTP management, and SSAS/SSIS restrictions with Data Factory.
Explore migration options to move on-premises databases to a managed instance. Choose between the migration service for minimal downtime, or backup and restore using blob storage and a restore command.
Compare general-purpose and business-critical service tiers for managed instances, highlighting storage and compute options, always-on availability, and the read-only replica for reporting to boost performance.
Explore management operations for Azure managed instances, including instant deployment, instant update, and instant deletion, with typical timelines and nuances for first versus subsequent deployments.
Provision a new managed instance, configure a virtual network and a jump-box virtual machine, connect with SQL Server Management Studio, create databases, and review cleanup to avoid costs.
Discover how Azure secures deployment options and databases with a layered defense—network security, access management, threat protection, and information protection, including firewall, directory integration, always encrypted, and transparent data encryption.
Secure Azure database by configuring network firewall rules, enforcing authentication and authorization with Azure Active Directory and SQL, enabling auditing and advanced threat protection, and applying encryption and data masking.
Deploy the azure managed instance into a virtual network to enable advanced security, secure on-premises connectivity via VPN, and private IP endpoints with dedicated compute and storage, in single-tenant infrastructure.
Explore data encryption across rest, in motion, and in use, covering symmetric and asymmetric encryption, key vault usage, and SQL Server always encrypted with deterministic and randomized options.
Apply dynamic data masking in SQL Server to protect sensitive data on the fly, using default, email, and credit card functions. Learn how portal and T-SQL implementations display masked results.
Explore high availability with multi-instance deployments and synchronous regional replication to prevent outages, and disaster recovery planning for unplanned site failures with RTO and RPO considerations.
Define recovery time objective and recovery point objective (RPO) for disaster recovery, and show how lower RTO and replication with standby regions reduce data loss.
Explore Azure SQL Database high availability and disaster recovery options across regions and within regions. Compare standard and premium models, read scale-out, and failover groups for seamless continuity.
Explore Azure SQL Database backup and restore essentials—full, differential, and transaction log backups, point-in-time and long-term retention, two-region geo-redundant storage, and in-portal restore workflows.
Learn to scale Azure SQL databases using vertical scale up and scale down, and horizontal scale out with read-only replicas and shards for tenancy and geo needs.
Explore portal monitoring for SQL databases, including activity logs, metrics, CPU alerts, and diagnostic settings, with destinations like Log Analytics, storage, and Event Hub.
Start with the portal to get quick answers using intelligence-backed analysis, then use dynamic management views, query store, and extended events to troubleshoot performance.
Explore automatic tuning and intelligent performance in a school database, including performance recommendations, index advisor, parametrized queries, and schema issue fixes, with automatic applying and verification.
Analyze resource consuming and long running queries using query performance insight tools, exploring cpu, data io, log io, execution counts, and the query store's capture and retention policies.
Explore cloud data warehousing with Edwardson Apps Analytics, covering MPP architecture, table types, partitioning, distribution, loading techniques, security layers, and backup and restore, through hands-on demos.
Explore cloud data warehousing as a fully managed, cost-efficient platform and see how Microsoft’s analytics service unifies ingestion, preparation, and serving with compute nodes and connectivity.
Explore how cloud data warehousing differs from on-premises, with platform as a service enabling on-demand scaling, cost control, and faster time to insight via MPP.
Explore how Synapse Analytics unifies data integration, storage, data warehousing, and big data analytics in a single workspace, enabling end-to-end analytics with serverless and dedicated pools.
Create a dedicated sql pool in Azure by provisioning a server, configuring firewall rules, and pausing compute to save costs, with deployment via demo inside Synapse Analytics service.
Learn to connect an Azure dedicated SQL pool to SQL Server Management Studio, configure firewall rules by adding your laptop IP, and access the AdventureWorks sample database for analytics.
Create and configure an Azure Synapse Analytics workspace, including subscription, resource group, region, data storage, and access roles, then explore Synapse Studio and endpoints.
Explore Synapse Studio v2 to ingest, develop, integrate, monitor, and visualize data within one interface, using pipelines, notebooks, data flows, and linked services for end-to-end analytics.
Create a dedicated SQL pool and an Apache Spark pool from within Synapse Studio. Learn the differences between serverless and dedicated pools, and how pausing saves costs.
Explore using a dedicated SQL pool to create and populate tables, run queries, and analyze 2.8 million New York taxi records with round robin and hash distributions.
Explore data from multiple sources with a spark notebook and pool; load the New York taxi parquet dataset into a dataframe, and analyze and move data across storage and databases.
Learn to analyze data with a serverless sql pool for ad hoc queries, explore blob storage and external sources, and create a serverless database to link external data.
Shows using data factory inside the SEINABO workspace to copy data from SQL Server to a data warehouse, configuring source and destination, creating a pipeline, and monitoring its run.
Explore how to monitor resources, activities, and queries in Synapse Analytics Studio using the monitor tab, including pipelines, integration runtimes, spark, serverless, and dataflow, with logs and query details.
Azure Synapse unites data warehousing, big data analytics, and data integration in a single unified service, enabling end-to-end analytics and cross-team collaboration from data to visualization.
Azure Synapse benefits highlights a unified analytics workspace for data preparation, management, warehousing, big data, and AI tasks, enabling scalable, secure data ingestion and dashboards.
Explains the Synapse unified experience for data engineers to connect to data sources, ingest, transform, and place data into storage from multiple sources, including streaming and transactional data.
Explore cloud warehousing benefits and the unified Snapp's analytics service, a third-generation data warehouse that accelerates time to market. Compare traditional warehousing with modern cloud architectures.
Explore how a massively parallel processing cloud data warehouse outperforms on-premises solutions while mastering storage, distribution (hash, round-robin, replicated), data types, and dimensional modeling using the Adventure Works case.
Explore the massive parallel processing architecture, where a control node coordinates parallel queries from compute nodes. Coordinate data transfers with data movement service, while storage remains separate from compute power.
Explore how a data warehouse uses distributions to store data and execute parallel queries across compute nodes. Compare hash, round-robin, and replicated sharding patterns for large versus small tables.
Explore data distribution methods in Azure data warehouse, avoid data skew, use hash distribution for joins, and round-robin for staging, while replicated tables fit small datasets with non-updatable distribution keys.
Choose the smallest data types and default lengths for integers and character columns, noting Unicode uses more space. Explore the three table types—clustered columnstore, heap, and clustered index—and their trade-offs.
Discover table partitioning in data warehouses, splitting a table into partitions by date to improve query performance, simplify loading and maintenance, and enable targeted queries with careful partition granularity.
Apply dimensional modeling to fact and dimension tables in a data warehouse, using hash distribution for facts and round-robin or hash for dimensions, with optional partitioning and clustering keys.
Analyze on-premises data warehouse distribution before migrating to azure, compare round-robin and hash distribution, and assess transaction history and product tables for efficient migration.
Explore the architecture of a data warehousing service, including storage and compute billing, scaling compute, data storage and distribution, three sharding patterns, and partitioning guidance for cloud migration.
Learn best practices for loading data in a data warehouse service using MBP architecture, compare SAS and Polybius loading methods, and explore Polybius setup with hands-on demos using Data Factory.
Explore best practices for parallel data loads in a data warehouse, including reader-writer pipelines, avoiding ordered data, and using temporary tables with round-robin distribution.
Compare single client and parallel loading methods, showing how the control node orchestrates queries while compute nodes scale, and how Polybius enables parallel loads from blob storage.
Compare loading with ssis and polybase, highlighting control node bottlenecks with ssis and the parallel, scalable loads from Azure blob storage via polybase, and outline key polybase setup steps.
SSIS loads a small dimension table from on-premises Adventure Works DW 2017 to the Adjure data warehouse using connection managers and one-to-one column mapping.
Export the on-premises table to a flat file, upload to blob storage, and load it into a data warehouse with PolyBase's six steps, then monitor and verify proper distribution.
Learn to load data from on premises to a data warehouse with Azure Data Factory, creating a data factory, linking blob storage, building a pipeline, and verifying the destination table.
Apply best practices for loading data to achieve fast times into a data warehouse, using control and compute node roles and load methods—single client and fully preload via Polybius—with demos.
This assignment guides you to move data from blob storage to a data lake and then into a synapse data pool using a data factory pipeline, outlining six steps.
Migrate data from a data lake to Synapse using data factory pipelines. Follow six steps—from creating storage and data factory accounts to moving files and validating the destination table.
Explore the five-layer defense to secure your adjure sequel database, monitoring traffic for suspicious patterns and enforcing network access, authentication, authorization, and data encryption.
Enable advanced data security to discover and classify sensitive data, assess vulnerabilities, and detect anomalous activity with data discovery and classification, vulnerability assessment, and advanced threat protection.
Auditing enables tracking of database events and audit logs to storage, log analytics, or event hub, letting you monitor activities and analyze history to identify threats and violations.
Learn how to secure Azure data warehouse by configuring firewall rules and virtual network rules to restrict access by IP addresses and subnets, with encrypted connections by default.
Learn how transparent data encryption protects data in transit with transport layer security and all newly created databases at rest encrypted by default with aes-256, with keys stored securely.
Guard sensitive data with dynamic data masking, hiding values from unauthorized users without changing the database, using default, partial, email, and random functions to control visibility.
Master access control through authentication and authorization for the data warehouse, including Active Directory and server logins, to protect data with role and column level security and granular permissions.
Explore five-layer defense in securing Azure data warehouses, covering data discovery, vulnerability assessment, threat detection, auditing, firewall and network settings, authentication versus authorization, role and column security, and encryption.
Navigate configuration settings in Microsoft Edwardson Snaps Analytics Service portal, back up and restore a data warehouse, manage workload, and monitor with activity alert metrics, diagnostic settings, and resource health.
Explore the data warehouse overview, learn to scale compute nodes, and view the activity logs. Configure tags, logs, security features, backups, and maintenance settings for the Adjure Cycle data warehouse.
Explore streaming analytics, load data via the portal, query the data warehouse with the portal editor, and build dashboards and reports with Power BI integration.
Learn how to back up a data warehouse using snapshots, perform incremental backups, create automatic and user-defined restore points, and execute point-in-time restores to recover or duplicate databases.
Scale an azure synapse data warehouse by adjusting compute units, pausing idle time to save costs, and automating scale with powershell, cli, or data factory scripts to manage concurrency.
Navigate the Microsoft pricing page for a data warehouse, learn region and currency options, hourly compute and storage costs, reserved savings, snapshot storage, and SLA basics.
Manage data warehouse workloads by classifying requests, assigning priority, and isolating resources to meet SLAs, optimize concurrency, and control query performance with workload groups and classifiers.
Explore monitoring options for your data warehouse, including query activity, alerts, and metrics dashboards. Learn to view query plans, diagnose settings, and pin charts to the dashboard.
Explore resource health to monitor data warehouse service status, diagnose issues, view remediation steps, and learn how to raise a Microsoft support ticket.
Delete resources in the Azure portal by removing individual items or the entire resource group, noting that deletion is irreversible and affects data factory, data warehouse, and blob storage.
Learn to configure and optimize a data warehouse, including backup, restore, cost management, and workload management concepts like classification, importance, and isolation, with monitoring via query activity and diagnostic settings.
Explore limits of traditional databases for big data, see how a data lake handles unstructured and structured data, and query files directly in a geo data storage and analytics demo.
Explore why traditional databases fail with exploding data and how a hyperscale, native-format repository can handle big data analytics across volume, variety, and velocity.
Store structured and unstructured data in a data lake at any scale and in raw form, then load immediately and transform later for analysis.
Understand how a data lake architecture and Hadoop complement each other, and how technologies like Apache Kafka and modern data warehouses enable storage, real-time data, and analytics.
Explore how data lake evolution blends Hadoop's HDFS heritage with Azure Blob Storage, from Gen1 to Gen2, to support scalable, cost-efficient big data storage and analytics.
Compare blob storage and data lake to understand their roles in big data analytics. Data lake, built on blob storage, enables Hadoop integration and analytics without moving data.
Discover hierarchical namespace in data lake storage, enabling true folders and efficient directory operations. Leverage it to boost analytics performance and enable fast move, rename, delete, with Hadoop integration.
Learn how to create an Azure storage account and enable the hierarchical namespace to convert it into Azure Data Lake Storage Gen2, built on blob storage, using the portal.
Explore a data lake gen2 storage account overview, including containers, files, tables, and queues, along with metrics, activity logs, access controls, and lifecycle management to optimize data transfer and security.
Explore data lake storage Gen2 features, including ABFS driver integration with Hadoop ecosystems for seamless data access. Learn about scalable, cost-efficient storage, security, and unstructured data challenges.
Explore tools to ingest data into Azure data lake Gen2 from diverse sources, and understand the data flow and ingestion options.
Demonstrates ingesting data into a data lake Gen2 from on-premises using portal and storage explorer, creating containers and directories, uploading file types, and selecting hot, cool, or archive access tiers.
Demonstrates copying data from on premises to a data lake account using the easy copy utility, including setting account name and key and performing a recursive copy.
Copy data from azure blob storage to data lake Gen2 using data factory, including creating connections, configuring a source file, and running a copy pipeline to move data.
Move structured data from a SQL Server database to Gentoo data lake Gen2 using Azure Data Factory, creating a SQL Server connection and transferring the customer table.
Demonstrates moving data from Amazon S3 to Azure Data Lake Gen2 using data factory by configuring source and destination, creating a bucket and folders, and deploying a copy pipeline.
Explore the data lifecycle around a data lake, from ingestion and exploration to cleaning, transforming, and loading data for analysis. Discover how to analyze and visualize results using cloud tools.
Explore how data lake architecture separates storage from compute, enabling on-demand processing with transient clusters while keeping data stored safely beyond traditional Hadoop's tightly coupled storage and compute.
Learn to connect a cloud data lake to a database, analyze data in place with Scala, Python, or SQL, and manage authentication and mounting with a service principal.
Provision Databricks, clusters and workbook by creating a database service workspace, selecting premium tier, and building an interactive cluster with notebooks.
Mount a data lake to Databricks DBFS by creating a service principal, configuring access with storage explorer, and mounting using a prepared notebook with client ID, secret, and directory ID.
Explore taxi trip data in a data lake by loading, examining, analyzing, cleaning, and transforming it with Spark. Save the results as parquet and reload to the data link.
Learn Azure storage security fundamentals, including keys, SAS, and storage access policies, dual active directory controls with RBAC and ACL, network restrictions, encryption in transit and at rest, threat protection.
Learn how storage account access keys work, including q1 and q2 and their connection strings, and why rotating keys is essential to prevent security risks in production.
Apply least privilege by using a shared access signature to grant scoped blob storage access for a defined time, with optional IP and https restrictions, revocable via stored access policy.
Azure Active Directory authentication enables identity as a service for storage access, using app registrations, service principals, and role-based access control to grant scoped access across subscriptions.
Master Azure role-based access control by applying role assignments across scopes from management groups to resources. Understand built-in roles such as owner, contributor, and reader.
Explore how ACLs grant read, write, and execute permissions to users, groups, and service principals on storage accounts, containers, and files. See how this complements RBAC.
Lock down your storage account with firewall and virtual network rules, restricting access to approved IP ranges and subnets, while enforcing authentication and authorization, and enabling trusted Microsoft services.
Enable secure transport to encrypt data in transit across on-premises, cloud storage, and regions; use blindsight encryption for double encryption and vault-managed keys.
Encryption at rest is enabled by default for all storage accounts, using 256-bit encryption with Microsoft managed keys or a custom Key Vault key, and cannot be disabled.
Enable advanced threat protection in the portal to secure storage and detect anomalous access, with alerts showing the account, IP, potential causes, and remediation guidance.
Explore how the Azure activity log records administrative changes, RBAC updates, key changes, and service health events across a subscription, enabling routing to storage, event hubs, or log analytics.
Explore how to view and filter activity logs for resources, inspect operation results and initiators, and export logs to log analytics or diagnostic settings for analysis.
Explore how metrics capture time-stamped values at regular intervals, including transaction latency and end-to-end latency, enable alerting on thresholds, and store data in a time-series database for near real-time monitoring.
Explore how to monitor resources with metrics in the adjure monitor service, focusing on storage accounts, capacity, transactions, latency, and the use of filtering, splitting, and alerting to troubleshoot performance.
Explore the new storage monitoring insights to analyze availability, transactions, and end-to-end vs server latency, using customizable workbooks to troubleshoot performance, failures, and capacity across blob, file, and table APIs.
Create and manage Azure alerts in the monitor service using storage account scope and metric or log signals to trigger actions via groups that include email, ITSM, and automation.
Explore diagnostic settings and locks, how events generate logs sent to log analytics, analyze them with a custom query language, and set alerts for issues.
Configure diagnostic settings for storage accounts to monitor with metrics and logging, with retention up to 365 days. Logs are stored as blobs in containers and accessible via storage explorer.
Improve data lake throughput by addressing ingestion bottlenecks, enabling parallel reads/writes, standardizing file sizes and naming, keeping services in the same region, and batching data to reduce latency and costs.
Explore how NoSQL overcomes traditional RDBMS limits by enabling horizontal scaling and flexible, schema-less JSON data management for unstructured data and rapidly changing requirements.
Compare sql and no sql databases: sql is relational with static schemas and tables; no sql is non-relational with dynamic schemas and collections, enabling vertical scaling vs horizontal scaling.
Master four NoSQL data models: key-value stores, document stores, column stores, and graph databases, and learn their structures, use cases, and performance characteristics.
Explore Microsoft Azure NoSQL offerings, including storage options blob, table, file, and queue storage, plus Cosmos DB as a multi-model service with multiple APIs, and differentiate data lake from storage.
Azure storage accounts and blob types, including page blobs, and learn how worldwide, secure, and scalable storage supports disaster recovery, encryption, private endpoints, and cost management.
Provide an Azure storage account in the portal by selecting subscription, resource group, unique name, and location; compare standard and premium options and general-purpose v2.
Learn how azure storage replication protects data with six options across primary and secondary regions, including locally redundant storage, zone redundant storage, geo redundant storage, and read-access variants.
Provision an Azure storage account, choose replication and hot, cool, or archive access tiers, and configure lifecycle and connectivity options for secure, cost-efficient data storage.
Store massive unstructured data with Azure blob storage, an object store for text, images, videos, and backups. Choose block, append, or page blobs for flexible uploads and long-term storage.
Manage data lifecycle in blob storage with hot, cool, and archive tiers, using lifecycle policies to automate transitions, save costs, and improve data quality.
Explore high availability and disaster recovery options for Azure storage, including locally redundant storage, zone redundant storage, zero redundant storage, read data access, and manual failover.
Explore how Cosmos DB evolved from a document database to solve global distribution and scalability, addressing relational database limits for worldwide data needs.
Explore Cosmos DB features, a globally distributed, multi-model, fully managed database as a service with automatic indexing, schema-agnostic design, and turnkey global distribution for high availability.
Explore Cosmos DB's multi-model APIs: sequel API, mongered API, Cassandre API, Table API, and Gremlin API, and learn migration paths, API differences, and graph and document data modeling.
Compare Azure Table Storage and Cosmos DB Table API, highlighting global distribution capabilities, automatic indexing, and throughput versus storage pricing.
Provision a new Azure Cosmos DB account using the Core SQL API, configure resource group and location, enable geo-distribution and multi-region writes, set networking and encryption, then review and create.
Understand how Cosmos DB uses databases, containers, and items, while grasping throughput, partition keys, and API terminology across SQL API, Cassandra API, MongoDB API, Gremlin API, and Table API.
Learn how to measure and optimize Cosmos DB throughput and latency using request units, understanding RU basics, predictable costs, and alerts to scale throughput.
Learn how Cosmos DB achieves horizontal scalability with unlimited storage and unlimited throughput by partitioning data across containers and multiple machines.
Understand how Cosmos DB partitions a container into logical partitions via a partition key, mapped to physical partitions across machines, with examples like city or customer ID.
Learn how Cosmos DB enables dedicated vs shared throughput, comparing database-level and container-level provisioning, how throughput is allocated to containers, and the impact on reserved versus shared capacity.
Explore how container throughput splits across logical partitions in Cosmos DB and how to avoid hot partitions by selecting partition keys that evenly distribute data and queries.
Compare single partition queries in Cosmos DB that fetch a user’s data from a logical partition with cross partition fan out queries that incur higher costs and should be avoided.
Explore composite keys in Cosmos DB by choosing partition keys that spread data across many small logical partitions, respecting 2 megabyte per document and 20 gigabyte per partition.
Select a high cardinality partition key at container creation to avoid hotspots and remember you cannot change it later.
Discover automatic indexing in Cosmos DB, where every item's properties are indexed by default with no schema, and inverted indexing keeps indexes in sync through updates.
Demonstrates creating a Cosmos DB database and container, adding and querying items with a partition key, throughput controls, and using the data explorer to modify and explore data.
Learn how Cosmos DB's time to live enables automatic deletion of documents after a configured interval. Set on/off and per-item expiry to manage data retention.
Explore Cosmos DB’s global distribution and multi-region replication to replicate data across data centers, enable read and write near users, reduce latency, and support disaster recovery and business continuity.
Enable Cosmos DB multi region writes for read and write across data centers with low latency. Explore conflict resolution options: last writer wins, merge procedures, and conflict feeds.
Configure manual or automatic failover in Cosmos DB across multi-region write-enabled centers to maintain disaster recovery and seamless failover.
Explore Cosmos DB five consistency levels—strong, bounded staleness, session, consistent prefix, and eventual—and learn to balance availability, latency, and data order across data centers.
Learn to create a Cosmos DB account using Azure CLI, set consistency levels and failover options, define location and resources, and create a database and container from example code.
Explore Cosmos DB pricing and scaling options, including manual vs autoscale throughput, multi-region and multi-master costs, and storage charges to optimize performance and budget.
Learn to monitor Cosmos DB costs and performance using Azure Monitor, assess partition key and container usage, view metrics and alerts, and optimize throughput and storage.
Explore how to monitor Cosmos DB costs within the Azure portal using metrics dashboards, alerts, and diagnostic settings, including throughput, storage, latency, replication, and logs.
Explore role-based access control, firewall and virtual networks, key management (primary and secondary keys, read-only keys), cross-origin resource sharing (CORS), private endpoints, and advanced security for Cosmos DB.
Explore high availability and disaster recovery in Cosmos DB, including local and global replication. Learn backup and restore, automatic failover, and multi-region failover in the portal.
Explore live data processing through the three core components—event producer, event processor, and event consumer—for real-time actions with Spark Streaming and Azure Stream Analytics.
Explore Azure streaming analytics and how it ingests, processes, and delivers real-time insights from fast-moving live data using Event Hubs and IoT Hub, with blob storage outputs.
Explore streaming analytics by grouping timestamped events into five-second buckets with tumbling, hoping, sliding, and session windows to compute averages, counts, min, and max.
Learn about tumbling windows that divide streaming data into non-overlapping 10-second buckets. Use group by to count events per bucket and read results for 0 to 10, 10 to 20, and 20 to 30 seconds.
Learn how hopping windows in streaming analytics create overlapping 10-second windows that move every five seconds to count tweets.
Learn how sliding window with a fixed 10-second length processes events on a timeline, starting at each new event, creating overlaps and counting events per window.
Analyze session windows in streaming analytics: a non-overlapping, event-started window that ends after five minutes of silence or at a ten-minute maximum.
Create blob storage input and output containers, configure an Azure streaming analytics job with input, output, and processing logic, and run it to process json files.
Learn how to ingest data from the Azure IoT Hub with a Stream Analytics job, route to blob storage, and filter sensor data by temperature above 27.
Monitor streaming analytics via the portal, dotnet sdk, or visual studio, track metrics like utilization, runtime errors, and watermark delay, and set alerts with diagnostic logs.
Optimize a stream analytics job by tuning input, output, and query processing; monitor CPU utilization, scale streaming units, and use partitioning to parallelize across input and output partitions.
Discover spark basics, including the in-memory resilient distributed dataset, lazy evaluation, and lineage. Compare spark with Hadoop, learn data frames and datasets, and explore APIs across Python, Java, and Scala.
Explore how Spark becomes easier with data breaks, as it manages infrastructure, autoscales clusters, and enables collaboration through an elegant UI.
Discover how Azure Databricks enables cloud data integration by ingesting diverse sources into a data lake via Azure Data Factory, then processing data for analytics-ready outputs.
Demonstrates connecting data breaks with data link and data lake to analyze and process data in place, using scala, python, or sql.
Provision an Azure Databricks workspace, configure a premium cluster, and create a notebook to run sample commands, illustrating workspace, cluster, and notebook workflows.
Mount a data lake to Databricks dbfs using a service principal and app registrations, configure credentials, and set read, write, and execute access.
Explore, analyze, clean, transform, and load taxi data in Databricks from the data lake. Learn to read csv, apply Spark transformations, and write parquet outputs back to the data lake.
Explore Azure Databricks clusters, including interactive notebooks vs automated job clusters, their scaling, auto-termination, and how standard and high concurrency modes affect performance and cost.
Explore Azure Databricks components such as workspaces, databases, tables, notebooks, and jobs, and learn to run multi-language notebooks, schedule jobs, manage clusters, and use libraries for end-to-end data workflows.
Explore databricks monitoring options, including the ganglia monitoring system embedded with the database for live metrics and logs, and gravagna monitor with setup steps.
Why Microsoft Azure Data Engineer?
Microsoft Azure Date Engineering is one of the fastest-growing and in-demand occupations among Data Science practitioners.
According to a 2019 Dice report, there was an 88% year-over-year growth in job postings for data engineers, which was the highest growth rate among all technology jobs.
If you are interested in this domain, there cannot be a better time to start learning these Microsoft Azure Data Solution technologies.
Expected Outcomes
After this course:
Microsoft Azure SQL Server - You will be able to identify the right Azure SQL Server deployment option, purchasing model and service tier according to requirements and will be able to deploy it in the cloud. You will be able to set up a security configuration to secure your database.
Microsoft Azure SQL Data Warehouse - You will able to deploy Azure Synapse Analytics (formerly known as Azure SQL Data warehouse) in Azure Cloud environment. You will have good internal MPP architecture understanding, and so you will be able to analyze your on-premises data warehouse and migrate data to Azure Data Warehouse.
Microsoft Azure Data Lake - You will be able to create Azure Data Lake storage account, populate it will data using different tools and analyze it using Databricks and HDInsight
Microsoft Azure Data Factory - You will understand Azure Data Factory's key components and advantages. You will be able to create, schedule and monitor simple pipelines.
Hadoop Basics - You will learn the fundamental understanding of the Hadoop Ecosystem and 3 main building blocks. This module will prepare you to start learning Big Data in Azure Cloud using HDInsight.
Microsoft Azure HDInsight - We will understand how HDInsight makes Hadoop easy, and we will go through simple demo where we’ll fetch data from Data Lake, process it through Hive and later will store data in SQL Server.
Cosmos DB - You will understand all basic concepts of Cosmos DB, and will be able to create a new database, container, documents, and choose right partition key, configure the global distribution and use other important features
Streaming Service - I will add details soon
Databricks - Video lessons added, I will add details here soon
Storage Service - Video lessons added, I will add details soon
Intended Audience
Beginners in Microsoft Azure Platform
Microsoft Azure Data Engineers
Microsoft Azure Data Scientist
Database and BI developers
Database Administrators
Data Analyst or similar profiles
On-Premises Database related profiles who want to learn how to implement these technologies in Azure Cloud.
Anyone who is looking forward to starting his career as an Azure Data Engineer.
Level
Beginners (100)
"if you are already experienced and working on these technologies, this may not be the best course for you."
I have a few crash courses (Free) for absolute beginners, you can find links on my website.
Prerequisites
Basic T-SQL and Database concepts
Microsoft Azure Free trial Subscription
Language
English
If you are not comfortable in English, please do not take a course, captions are not good enough to understand the course.
What's inside
Video lectures, PPTs, Demo Resources, Quiz, Assignment, other important links
Full lifetime access with all future updates
Certificate of course completion
30-Day Money-Back Guarantee
Course updates
I am consistently adding new content in terms of both lectures and demos. Please find below logs for past and upcoming changes.
Previous Updates:
Jan 2021 - Revised Synapse module with all new changes
Aug 2020 - Data Encryption related lessons added
July 2020 - Monitoring and Optimization of services
July 2020 - AMS (Azure Monitoring Service)
July 2020 - Re-recorded Azure Data Factory with all new updates
July 2020 - NoSQL Databases in Microsoft Azure (Azure Blob Storage)
July 2020 - Added Azure Databricks Module
July 2020 - Added Azure Streaming Analytics Module
June 2020 - Added Azure Cosmos DB Module
May 2020 - Full-length course on Data Lake
Apr 2020 - Added modules on Azure cloud computing basics, and Data Engineer profile
Apr 2020 - Added crash course on cloud computing and data warehouse
Mar 2020 - Added Azure Synapse Analytics
Feb 2020 - Added Azure SQL Database
Oct 2019 - Added Azure HDInsight
Sep 2019 - Added Azure Data Factory
Aug 2019 - Course Created on Azure Data Lake
Course In Detail
By the end of this course
Microsoft Azure SQL Server -
You will know the advantages of Azure Database over on-premises Database
You will be able to provision all three types of Azure Database PaaS Deployments (Single, Elastic Pool, Managed Instance)
You will learn different purchasing models and Service tiers and how to configure them according to requirements.
You will be able to provision Azure Database in Virtual Machine (IaaS)
You will be able to deploy Elastic Pool Database and able to add/remove more databases from the pool.
You will be able to deploy the Managed Instance Database, configure it and connect it with your SQL Server Management studio.
Microsoft Azure Synapse Analytics Service
You will learn why we should consider warehousing solutions in the cloud.
You will learn about Microsoft's brand new Azure Synapse analytics service, and how this service brings together enterprise data warehousing and Big Data analytics, and provide a unified experience to ingest, prepare, manage, and serve data for immediate BI and machine learning needs.
You will learn the difference between Traditional vs Modern vs Synapse Architecture.
You will learn Azure famous MPP (Massive Parallel Processing) architecture
You will also learn many internal concepts like Data Distribution, Sharding, and partitioning.
You will learn different migration methods and best practices.
You will learn PolyBase setup (Best Loading method)
You will perform a lot of demos to migrate data to Azure SQL Data warehouse using SSIS, Data Factory and PolyBase.
Microsoft Azure Data Lake Gen1
You will learn the limitations of traditional database systems to handle the Big Data revolution.
You will learn the difference between Azure Data Lake, SSIS, Hadoop and Data Warehouse.
You will be able to create Azure Data Lake Gen1 storage account, populate it will data and analyze it using U-SQL Language.
Microsoft Azure Data Factory -
You will understand Azure Data Factory's key components and advantages.
You will be able to create, schedule and monitor simple pipelines.
Hadoop Basics
You will learn a fundamental understanding of the Hadoop Ecosystem and 3 main building blocks.
This module will prepare you to start learning Big Data in Azure Cloud using HDInsight.
Microsoft HDInsight -
You will learn what are the challenges with Hadoop and how HDInsight solves these challenges.
You will also learn Cluster types, HDInsight Architecture and other important aspects of Azure HDInsight.
You will also go through demo where we’ll fetch data from Data Lake, process it through Hive and later will store data in SQL Server.
Some students Feedback
One of the most amazing courses i have ever taken on Udemy. Please don't hesitate to take this course. The instructor is really professional and has a great experience about the subject of the course. - Khadija Badary
Very nicely explained most of the concepts. a must have course for beginners - Manoranjan Swain
I appreciate this course explaining everything in great detail for a beginner. This will assist me in overcoming challenges at my work - Benjamin Curtis
Good course for Beginners. Labs are really helpful to grasp the concept. Thank you - Sapna