Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Getting Started with Big Data and the Hadoop Ecosystem
Rating: 3.6 out of 5(36 ratings)
1,236 students

Getting Started with Big Data and the Hadoop Ecosystem

Deep dive into the world of Big Data and Hadoop! Learn the need for Big Data analysis, technologies and distributions
Last updated 12/2016
English
English [Auto],

What you'll learn

  • Understand what Big Data is and its history
  • The need for Big Data in today's information world
  • Fundamentals of Hadoop
  • Apache
  • The most popular Hadoop Distriubters inlcuidng Amazon EMR, Cloudera, HDInsight, MapR, and Hortonworks

Course content

8 sections64 lectures11h 4m total length
  • Introduction to Hadoop Part 111:30

    Define Hadoop as a Java-based framework for processing large data sets in distributed computing, outline core components like HDFS and MapReduce, and introduce clusters, history, ecosystem, versioning, and distributions.

  • Introduction to Hadoop Part 211:29

    Explore the Hadoop ecosystem for distributed storage and processing on commodity hardware, with MapReduce (map, shuffle, reduce) and a focus on scalability, agility, and loading data without a fixed schema.

  • Hadoop History and Background Part 111:53

    Trace Hadoop’s origins from a 2002 search engine project to an open source Apache platform powered by commodity hardware. Learn how Google's file system and MapReduce papers shaped its birth.

  • Hadoop History and Background Part 214:36

    Trace how Yahoo helped Hadoop's rise, then how Hortonworks and Cloudera formed, and examine the open source ecosystem, distributions, and fragmentation in the Hadoop market.

  • What Makes Hadoop Special Part 18:37

    Hadoop delivers cost effectiveness by running on commodity hardware and a pay-as-you-go cloud. It runs on Linux, Mac OS X, or Windows with Java, offering fault tolerance and scalable nodes.

  • What Makes Hadoop Special Part 28:30

    Discover how Hadoop's distributed file system enables reliable, scalable storage with fault-tolerant replication, speculative execution, and MapReduce-driven computation across nodes.

  • Hadoop Ecosystem Part 110:03

    Hadoop is a framework, not an application, combining DFS (dubious file system) and MapReduce to store and process data, with requests flowing through MapReduce before storage.

  • Hadoop Ecosystem Part 29:35

    Explore Hive, a data warehouse on top of Hadoop, using HiveQL to run queries against tables stored in distributed file system, and compare it with Pig, data flows scripting language.

  • Hadoop Ecosystem Part 311:31

    Explore the Hadoop ecosystem core components—hdfs, mapreduce, hive, and pig—and learn how scoop writes structured data and flume handles unstructured streaming data, with base enabling real-time analytics.

  • Versions in Hadoop Part 19:37

    Explore how Hadoop versioning works, including major, minor, and point releases, and how MapReduce v1 and v2 map to Hadoop 1.x and 2.x, with guidance on upgrading and release compatibility.

  • Versions in Hadoop Part 27:56

    Explore how Cloudera distributions package Apache Hadoop, detailing CDH versions from 3 onward, MRV and MRV2 mapreduce engines, Yarn, and the Cloudera Manager options, with notes on CDH 5.4.

  • Popular Hadoop Distributions Part 111:16

    Explore popular Hadoop distributions such as Cloudera, Amazon EMR, Hortonworks, MapR, and Azure, and understand how open source Apache Hadoop becomes vendor packaged solutions with enterprise data hub features.

  • Popular Hadoop Distributions Part 210:57

    Leverage Amazon EMR’s elastic, cost-efficient data processing with multi-cluster provisioning and on-demand resizing. Integrate with S3, HDFS, and DynamoDB via EMRFS for flexible storage and processing.

  • Popular Hadoop Distributions Part 38:54

    Explore major Hadoop distributions and their integration with Amazon EMR and data pipelines. Compare Hortonworks and MapR offerings, including platform features and partnerships.

  • Popular Hadoop Distributions Part 48:58

    Explore popular Hadoop distributions, including Windows Azure hd inside, and examine how they deploy clusters with Azure blob storage, dfs, and polybase.

Requirements

  • None. This course is perfect for beginners and non-technical people.

Description

Ever curious what the term Big Data is? What's Hadoop and what's up with the Elephant? This course is for anyone who wants a clear structured introduction to the world of Big Data and Hadoop which has been labelled as the next generation platform for data processing because of its low cost and ultimate scalable data processing capabilities. 

Who's it for?

This course is for anyone who works with data analytics; including developers, analysts, marketers, or anyone generally interested in the topic.

What will I learn?

This course starts out by giving you the history of Big Data and why we need Big Data analysis in today's data driven world.  We then go into the history of Hadoop, how it was conceived, and the technology behind it. 


The latter half of this course will go into the top Hadoop Distributors. These are the companies taking the open source framework of hadoop and creating innovative products and solutions to meet the demands for Big Data technology. The companies covered are Amazon, Cloudera, HDInsight, MapR and Hortonworks.

Who this course is for:

  • Anyone interested in a foundational understanding of Big Data & Hadoop