
Explore how Hadoop enables big data analysis, tackle common implementation challenges, and see a live demo that fetches data, processes it with Hive in an HDInsight cluster, and stores results.
Learn how to create a free Azure subscription, claim a $200 credit for 30 days, and explore always-free services like blob storage, sql database, and functions.
Explore how the Azure portal is the primary interface to create, manage, and monitor resources and services. Customize dashboards, use search, switch subscriptions, and access cloud shell for streamlined workflows.
Explore how Azure offers cloud services across categories like AI and machine learning, analytics, compute, database, container, and storage; everything is a service in the cloud, with usage-based pricing.
Explore Azure resource management by organizing resources with management groups, subscriptions, and resource groups, and apply policies, budgets, and access controls.
Create and manage resource groups within Azure by linking subscriptions and management groups. Assign a metadata region, allow cross-region resources, and control access, costs, and lifecycle with locks and policies.
Tag resources in Microsoft Azure by adding metadata through tags on resources, resource groups, and subscriptions, then use tags to organize, search, and manage cost and billing.
Learn how to delete resources and resource groups to control costs, set budgets, and receive email alerts through cost analysis and forecast.
Discover the essentials of Hadoop, including the Hadoop distributed file system, MapReduce, and YARN, and why these components address data needs traditional systems can't meet.
Discover why distributed computing is essential for big data and how Hadoop enables scalable storage and processing as Facebook and Google data grows.
Compare monolithic single-server systems with distributed clusters of commodity machines, where data is partitioned across nodes and scaled linearly by adding more machines, coordinated by fault-tolerant cluster software.
Learn how distributed systems coordinate thousands of machines using the Google file system and MapReduce, and how Hadoop uses DFS, MapReduce, and Yarn to store and process data.
Compare Hadoop and rdbms by data structure. Hadoop uses schema-on-read for unstructured data, while rdbms relies on fixed schemas and strong transactions.
Gain a concise overview of Hadoop within a distributed computing framework, understanding why distributed systems outperform a single computer and the role and purpose of Hadoop in this ecosystem.
Explains why implementing on-premise Hadoop for big data projects is hard due to upfront hardware costs, scalability limits, and a shortage of skilled data engineers and Hadoop admins.
Simplifies Hadoop with a cloud-based, fully managed HDInsight service that eliminates upfront hardware costs and enables elastic, scalable computing with pay only what you use.
Examine HDInsight's integration with Hortonworks, Microsoft's use of Apache projects such as Pig, Spark, Hive, Storm, and Kafka, and how security, Active Directory, and log analytics support big data workloads.
Explore Azure HDInsight cluster types, including Hadoop, Spark, Kafka, HBase, and Storm, to tailor scalable, cost-efficient big data processing with decoupled architectures and persistent blob storage.
Understand the high level HDInsight architecture with a resource manager and a node manager, plus decoupled storage using blob storage for dual storage in HDInsight.
Explore a demo of Hadoop and Azure HDInsight basics by performing a batch processing workflow: upload a file, process with Hive SQL, and export results to SQL Server using scoop.
Create an Azure Data Lake Storage Gen2 source and a SQL Server destination by setting up the storage and database, then configure connectivity, firewall, and flight delay data output table.
Enable secure service-to-service authentication with managed identities in Azure, avoiding exposed credentials; assign system-assigned or user-assigned identities for access to storage and Azure Data Factory.
Create a managed identity to represent the inside cluster, then grant it role-based access to the Gentoo storage account and SQL Server, enabling secure interlinking of storage and databases.
Learn to create an Azure HDInsight interactive query cluster, configure memory-optimized settings, enable Hive SQL on Hadoop, set up Gen 2 data lake storage, and secure credentials for high availability.
Ambari overview and UI explains how Ambari, the open-source Hadoop management platform, handles cluster administration, monitoring, and configuration through a graphical dashboard with alerts, metrics, and service control.
Ingest a CSV dataset into data lake storage on Azure HDInsight, then run interactive Hive queries to create a new table, using browser-based tools and Storage Explorer for file upload.
learn how to use Hive's interactive query editor to create external tables, ingest csv data into delays_raw and delays, and fetch data from the data lake storage with mapreduce-backed queries.
Learn how Hive on Hadoop converts unstructured data into tabular form, creates a city delay list with average delay time, and saves it to an output folder for loading.
Authenticate to a SQL server from Sqoop, configure firewall rules, and export data from Gen 2 storage into the SQL database.
Demonstrate Hadoop basics in Azure HDInsight, compare command line and graphical interfaces, and show how to integrate a cluster with storage, SSL, and authentication.
Expected Outcomes
Hadoop Basics
You will learn a fundamental understanding of the Hadoop Ecosystem and 3 main building blocks.
This module will prepare you to start learning Big Data in Azure Cloud using HDInsight.
Microsoft Azure HDInsight
You will learn what are the challenges with Hadoop and how HDInsight solves these challenges.
You will also learn Cluster types, HDInsight Architecture, and other important aspects of Azure HDInsight.
You will also go through demo where we’ll fetch data from Data Lake, process it through Hive, and later will store data in SQL Server.
We will understand how HDInsight makes Hadoop easy, and we will go through simple demo where we’ll fetch data from Data Lake, process it through Hive, and later will store data in SQL Server.
Intended Audience
Beginners in Microsoft Azure Platform
Microsoft Azure Data Engineers
Microsoft Azure Data Scientist
Database and BI developers
Database Administrators
Data Analyst or similar profiles
On-Premises Database related profiles who want to learn how to implement Hadoop components in Azure Cloud.
Anyone who is looking forward to starting his career as an Azure Data Engineer.
Level
Beginners (100)
"if you are already experienced and working on these technologies, this may not be the best course for you."
Prerequisites
Basic T-SQL and Database concepts
Language
English
If you are not comfortable in English, please do not take a course, captions are not good enough to understand the course.
What's inside
Video lectures, PPTs, Demo Resources, Quiz, Assignment, other important links
Full lifetime access with all future updates
Certificate of course completion
30-Day Money-Back Guarantee
Some students Feedback
One of the most amazing courses i have ever taken on Udemy. Please don't hesitate to take this course. The instructor is really professional and has a great experience about the subject of the course. - Khadija Badary
Very nicely explained most of the concepts. a must have course for beginners - Manoranjan Swain
I appreciate this course explaining everything in great detail for a beginner. This will assist me in overcoming challenges at my work - Benjamin Curtis
Good course for Beginners. Labs are really helpful to grasp the concept. Thank you - Sapna