
Introduction
3Vs of Big Data
Learn how big data workloads require scale-out, distributed infrastructure to process large volumes across many inexpensive servers, leveraging machine learning and natural language processing.
| New forms |
| New languages |
| New hardware |
Explore the benefits of big data, including timely, accessible data that simplifies finding and managing information, with web analytics, social analytics, CRM, and e-mail marketing.
Discover how big data boosts data relevance and security, reduces costs with cloud-based analytics, augments existing architectures, and enables faster decisions through in-memory analytics, driving new products, services, and revenue.
Explore how increasing storage capacity meets growing data needs, address access speed, and enable replication for availability, plus map‑reduce style analysis combining data from multiple disks.
Explore how to query big data with map-based processing, examining brute-force vs selective processing, ad hoc queries, and analyzing geographic distribution across hundreds of gigabytes.
| HPC |
| Grid computing |
Explore big data by querying a public dataset, download samples, and perform data processing and exploration to derive insights.
Learn to select names, define a destination, and work with a tree while filtering by gender and handling an error message.
Create and configure a big table instance within a big data Hadoop environment. Examine table design, storage choices, and performance considerations.
Explore the pub-sub topic, define topics and subscriptions, and recap how subscriptions relate to topic definitions, as discussed around examples like museum's collection.
Explore how Hadoop enables open source distributed processing across a cluster of computers, combining storage and processing to handle large data sets.
| Key features |
| Ecosystem |
| Vendor Integration |
| Introduction |
| File system |
| Hard disk space |
| Streaming access pattern |
| Nodes |
| Architecture |
| Name Node |
| Data Node |
| Job Tracker |
| Task Tracker |
| Client |
| Client Interaction |
| Client data distribution |
| Heart beat |
| Job Tracker |
| Job splits |
| Job tracker |
| Query process |
| Creation of new file |
| Data pipelining |
Analyze how rack descriptions demonstrate data replication across locations to improve availability and resilience, examining three copies, placement strategies, and disaster scenarios in big data environments.
Explore how hdfs write operation works: a client contacts the NameNode, checks file existence and permissions, then writes data with three replicas and acknowledges back to the client.
Explore selection of data nodes and how node distance influences block placement and replication, including multiple copies and processing time considerations.
Learn how serialization converts structured objects to a byte stream for network transmission and how deserialization restores them for inter-process communication, enabling communication between programs written in different languages.
| Block size |
| Command |
Explore HDFS caching of frequently accessed blocks in memory to boost performance, and explain fencing and failover using ZooKeeper-controlled standby name nodes to ensure a single active node.
This lecture explains HDFS federation as a feature to balance load by adding multiple name nodes that partition namespace and block pools, improving scalability and fault tolerance.
Explore high availability in HDFS by coordinating primary and secondary name nodes, standby name nodes, and fencing to prevent a single point of failure and data corruption.
| HAR Files |
| Commands |
Explore Hadoop 2.0 features, including federation with multiple namespaces and name nodes, high availability, and enhanced processing with binary compatibility for MapReduce and snapshots.
| Processor |
| Memory |
| Cluster size |
| Master node |
| Configurations |
Set up a Hadoop cluster on Unix by installing Java, creating dedicated Unix users, configuring HDFS, and using start-stop scripts to manage cluster operations and optional management tools.
| Data integrity |
| Block scan |
| Capture data |
| Organize |
| Integrate |
| Analyze |
| Act |
| Data |
| Capture |
| Analyze |
| Cloud computing |
| Big data implementation |
Explore using the Hue HDFS file browser to open existing files, upload and create new files, read contents, and download items while managing cache for efficient access.
Explore using the HDFS file browser to open directories, manage user group permissions, upload files in multiple formats, and observe file metadata while refreshing views for up-to-date listings.
Learn to create and manage a cloud SQL instance, reading cloud resources and applying a serious, practical approach to building cloud skills within the Big Data Hadoop course.
Learn to create a bucket in Google storage, ensure unique bucket names, and upload files and objects like photos and documents into the bucket.
| Introduction |
| Map |
| Reduce |
| Code |
| Architecture |
| Map Job |
| Splits |
| Map phase |
| Shuffle & Sort |
| Reduce phase |
| Key value pair |
| Map Function |
| Job Tracker |
| Job queue |
| Job tracker |
| Task tracker |
Task Distribution
| Key Value pair |
| Anatomy |
| Custom data types |
| Avro |
| SPOF |
| High availabilty |
Explore how to submit a job on the platform, discuss managing objects and segments, and understand why a real job may take time.
Explore the function of job design, viewing jobs as configurable maps with properties. Learn to create and manage jobs, use configuration types, and track statuses to reduce burden.
examine how a Hue metastore manager coordinates store data, including customers and employees, with real-time creation times, locations, and storage permissions to manage databases and data sets.
| Introduction |
| Layers |
Explore how YARN manages resources across a Hadoop cluster using a single resource manager and multiple node managers, launching containers to run application processes within constrained resources.
Hive provides a processing layer over data infrastructure to process data and store schema; not a relational database, open source by the Apache Software Foundation, used by Facebook and Amazon.
Discover Hive basics for data warehousing on Hadoop, enabling analysis of large data volumes by storing data in tables and databases and querying via a command line interface.
| Components |
| Hcatalog |
| Web H cat |
Discover Hive architecture by examining the file system, processing, and execution engine, then learn how interfaces and a database store schemas and data types with mappings.
Master Hive queries for big data analysis in the Hadoop ecosystem, learning to write and run selects, joins, and sorts, and to view results via dashboards and editors.
| Introduction |
| Components |
| Modes |
| Execution |
Explore a comparative overview of Pig, Hive, and MapReduce, highlighting language types, levels of abstraction, performance, and suitability for structured and unstructured data.
According to Forbes Big Data & Hadoop Market is expected to reach $99.31B by 2022 growing at a CAGR of 42.1% from 2015. McKinsey predicts that by 2018 there will be a shortage of 1.5M data experts. According to Indeed Salary Data, the Average salary of Big Data Hadoop Developers is $135k
This Big Data Hadoop - The Complete Course covers the topics from the basic level of beginner till the advanced professional levels required for Big Data Hadoop Certification. It's a must have course for prospective Big Data experts. The course covers Hadoop, HDFS, Map Reduce, YARN, Apache Hive, PIG, Impala, Scoop and ZooKeeper
Come join this course for a secure career!