
Define Hadoop as a Java-based framework for processing large data sets in distributed computing, outline core components like HDFS and MapReduce, and introduce clusters, history, ecosystem, versioning, and distributions.
Explore the Hadoop ecosystem for distributed storage and processing on commodity hardware, with MapReduce (map, shuffle, reduce) and a focus on scalability, agility, and loading data without a fixed schema.
Trace Hadoop’s origins from a 2002 search engine project to an open source Apache platform powered by commodity hardware. Learn how Google's file system and MapReduce papers shaped its birth.
Trace how Yahoo helped Hadoop's rise, then how Hortonworks and Cloudera formed, and examine the open source ecosystem, distributions, and fragmentation in the Hadoop market.
Hadoop delivers cost effectiveness by running on commodity hardware and a pay-as-you-go cloud. It runs on Linux, Mac OS X, or Windows with Java, offering fault tolerance and scalable nodes.
Discover how Hadoop's distributed file system enables reliable, scalable storage with fault-tolerant replication, speculative execution, and MapReduce-driven computation across nodes.
Hadoop is a framework, not an application, combining DFS (dubious file system) and MapReduce to store and process data, with requests flowing through MapReduce before storage.
Explore Hive, a data warehouse on top of Hadoop, using HiveQL to run queries against tables stored in distributed file system, and compare it with Pig, data flows scripting language.
Explore the Hadoop ecosystem core components—hdfs, mapreduce, hive, and pig—and learn how scoop writes structured data and flume handles unstructured streaming data, with base enabling real-time analytics.
Explore how Hadoop versioning works, including major, minor, and point releases, and how MapReduce v1 and v2 map to Hadoop 1.x and 2.x, with guidance on upgrading and release compatibility.
Explore how Cloudera distributions package Apache Hadoop, detailing CDH versions from 3 onward, MRV and MRV2 mapreduce engines, Yarn, and the Cloudera Manager options, with notes on CDH 5.4.
Explore popular Hadoop distributions such as Cloudera, Amazon EMR, Hortonworks, MapR, and Azure, and understand how open source Apache Hadoop becomes vendor packaged solutions with enterprise data hub features.
Leverage Amazon EMR’s elastic, cost-efficient data processing with multi-cluster provisioning and on-demand resizing. Integrate with S3, HDFS, and DynamoDB via EMRFS for flexible storage and processing.
Explore major Hadoop distributions and their integration with Amazon EMR and data pipelines. Compare Hortonworks and MapR offerings, including platform features and partnerships.
Explore popular Hadoop distributions, including Windows Azure hd inside, and examine how they deploy clusters with Azure blob storage, dfs, and polybase.
Big data describes large data sets that are structured, unstructured, or semi-structured and grow rapidly, challenging traditional databases. Devices, social networks, and sensor networks fuel this revolution.
Explore the three Vs of big data—volume, velocity, and variety—alongside real-world growth trends and storage scales from megabytes to exabytes.
Explore how fast-growing social media data drives the need for big data solutions, and how Hadoop enables scalable, parallel analysis of terabytes to petabytes.
Hadoop is a Java-based framework built on MapReduce and the Google File System, designed to process massive data sets on commodity hardware with parallel processing.
Identify big data using the 3 v's—volume, velocity, and variety—to classify data, distinguish transactional from behavioral data, and determine when Hadoop supplements relational databases.
Identify big data using the three v's—volume, velocity, and variety—by analyzing data in motion, velocity dimensions, and data utility over time, including unstructured formats and rdbms-hadoop contrasts.
Explore the limitations of traditional RDBMS in the big data era, examine data growth and the volume, velocity, and variety of data, and contrast with Hadoop and big data challenges.
Analyze the bottlenecks of traditional databases handling rapid data growth and RAM limits. Compare how Hadoop and distributed architectures improve scalability, addressing storage, read, and compute challenges.
Explore how traditional RDBMS and Hadoop compare, covering tables, keys, normalization, and joins, and explain scalability limits and the shift to distributed computing for big data.
Explore big data challenges in today's relational databases landscape and see how Hadoop's etche dfs and map reduce enable scalable, cost effective storage and parallel processing across thousands of nodes.
Explore how traditional rdbms face failures, consistency, and scalability challenges, and discover why Hadoop—handling volume, velocity, and variety—offers fault-tolerant batch processing for big data.
Discover the history and architecture of Apache Hadoop, including HFS and mass produce, and how they enable distributed storage and processing on commodity hardware with yarn.
Discover the Apache Hadoop architecture, including the Hadoop common package, HFS, name node and data node roles, job tracker and task tracker, YARN, and the MapReduce engine.
Explore the core components of Apache Hadoop, including the distributed file system and map reduce, with name node, data nodes, and job tracker enabling scalable, highly available processing.
Install and configure a single-node Hadoop cluster using the Hadoop distribution, MapReduce and HDFS, with prerequisites like Java and SSH, and prepare for YARN in the next class.
Explore yarn, the next generation mapreduce, and how the resource manager and application master allocate resources, run apps, and monitor tasks, plus hdfs with name node and data nodes.
Explore Amazon EMR and the Hadoop ecosystem, including data stores such as S3, HDFS, and DynamoDB, EMRFS, encryption, and tools like Hive, Pig, and Impala.
Explore the Hadoop ecosystem, including Spark for in-memory fault-tolerant processing, DAG-based data transformations, and Presto for interactive SQL queries, plus Amazon EMR and S3 integration for scalable deployment.
Discover Amazon EMR architecture, including data stores such as Amazon S3, HDFS, and DynamoDB, plus Hive, Pig, HBase, Impala, Spark, Presto, and deployment and security features.
Explore Amazon DynamoDB’s document and key-value data models, seamless scaling, high availability, streams, cross-region replication, triggers, atomic counters, strong consistency, and EMR/Redshift/data pipeline integrations.
Explore Amazon Redshift, a fast, fully managed data warehouse that scales from small to large, offering pay-by-byte pricing, columnar storage, and an MPP architecture for business intelligence tools.
Explore cloud storage options in the Hadoop ecosystem with S3, EBS, and Glacier, covering cross-region replication, versioning, encryption, lifecycle management, and snapshots.
Explore Amazon cloud storage options, including EBS durability, encryption, Glacier archives, vaults, access management, and data transfer services for secure, scalable big data workflows.
Trace Cloudera’s history from its 2008 founding in Palo Alto to the CDH distribution, Cloudera Manager, and the lineup of express, enterprise editions in the Hadoop ecosystem.
Explore Cloudera CDH, an open-source Hadoop distribution, delivering core Hadoop with HFS, MapReduce, and Impala for unified batch processing, SQL, and interactive analytics in the data hub.
Discover CDH tools like Cloudera search, Spark, and HBase to enable fast, scalable data exploration, real-time streaming, and unified analytics in the enterprise data hub.
Explore CDH tools like Apache Sentry for fine-grained access control and secure data, plus Kafka for real-time streaming and Akam Willow for scalable data storage.
Explore Cloudera Express, the Hadoop distribution that bundles CDH with Cloudera Manager, enabling automated deployment, centralized administration, monitoring, and diagnostics.
Explore Cloudera Enterprise, a paid Hadoop distribution with CDH and management. Discover how Cloudera Manager and Navigator support governance, performance, and analytics with HBase, Impala, and Spark.
Explore Cloudera Director's self-service, cloud-centric management for deploying and scaling CDH and Cloudera Enterprise clusters. Understand hybrid deployments, security and governance, and open standards in the Hadoop ecosystem.
Discover how HDInsight deploys and provisions Apache Hadoop clusters in the cloud. Learn the Hadoop ecosystem's core components, from HFS and MapReduce to Hive, Pig, Sqoop, Tez, and Zookeeper.
Discover how HBase on Azure HDInsight stores data in blob storage, enables real-time queries with Phoenix, and supports key-value and table-based big data workflows.
Explore how Apache storm on HD insight enables real-time analytics with scalable fault-tolerant topologies, Nimbus and worker supervision, and cross-language thrift support.
Explore how PolyBase brings the relational world to the Hadoop cloud, enabling external data sources, external tables, and external file formats to deliver fast, hybrid queries.
Explore how to set up an Azure storage account, provision an HDInsight Spark cluster, and run Spark SQL queries with Zeppelin and Jupyter notebooks to analyze sample data.
Trace MapR's history, from its 2009 founding. Explore MapR's distribution and editions, including standard, enterprise database edition, and enterprise, with MapR D-B NoSQL for online and analytical processing.
Explore MapR distributions for Hadoop, including M3 standard, M5 enterprise, and M7 enterprise database editions, with features like backward compatibility, broad project support, Kerberos and LDAP security, and high availability.
Learn security and data governance in the Hadoop ecosystem with MapR, covering extensible authentication, granular access controls, auditing, encryption, and data lineage with volumes and snapshots.
Explore MapR add-ons, including Apache Drill, Apache Sparc, and Apache Soler, with commercial support options, enabling rapid, governance-friendly, self-service analytics on big data.
Trace the history of Hortonworks, its open leadership, and the Hortonworks data platform (HTP) built on Apache Hadoop, featuring HFS and yarn for enterprise on premise and cloud deployments.
Discover the HDP platform from Hortonworks, covering data management, data access, governance, security, and operations, with Hadoop components such as HDFS, YARN, Hive, Solr, Storm, and Spark.
Explore the Hadoop ecosystem with HDFS and YARN, learning how scalable, fault-tolerant storage and resource management power batch, interactive, and real-time data workloads, plus security, upgrade, and high availability features.
Explore how HDFS uses a name node, data nodes, and 128 MB blocks replicated across the cluster. See how YARN enables multi-tenancy, resource management, and scalable scheduling for Hadoop.
Load and manage data according to policy within the Hortonworks open enterprise platform, leveraging Hadoop, HDFS, and YARN for secure authentication, authorization, and data protection across on-premise and cloud deployments.
Explore authentication, authorization, and data protection in Hadoop using Apache Knox gateway and Apache Ranger; simplify Kerberos, enable central security policies, audits, and enterprise-ready access control.
Explore deploying Hortonworks data platform in the cloud, including Windows and Azure options, with Microsoft and Red Hat partnerships enabling scalable, enterprise Hadoop across on-premise and cloud environments.
Discover how Apache Hadoop in the cloud integrates with Hortonworks, Red Hat Storage, and OpenJDK. Leverage Rackspace private cloud for elastic, secure on-demand analytics.
Ever curious what the term Big Data is? What's Hadoop and what's up with the Elephant? This course is for anyone who wants a clear structured introduction to the world of Big Data and Hadoop which has been labelled as the next generation platform for data processing because of its low cost and ultimate scalable data processing capabilities.
Who's it for?
This course is for anyone who works with data analytics; including developers, analysts, marketers, or anyone generally interested in the topic.
What will I learn?
This course starts out by giving you the history of Big Data and why we need Big Data analysis in today's data driven world. We then go into the history of Hadoop, how it was conceived, and the technology behind it.
The latter half of this course will go into the top Hadoop Distributors. These are the companies taking the open source framework of hadoop and creating innovative products and solutions to meet the demands for Big Data technology. The companies covered are Amazon, Cloudera, HDInsight, MapR and Hortonworks.