Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Big Data and Hadoop : Interactive Intense Course
Rating: 3.9 out of 5(23 ratings)
1,808 students

Big Data and Hadoop : Interactive Intense Course

Managing Big data using Hadoop tools like MapReduce, Hive, Pig, hBase and many more and Able to crack Cloudera CCA 175.
Last updated 10/2018
English

What you'll learn

  • In late 1990s, Developers and Programmers were generating data through coding, in late 2000s everyone on Social media generating data on FB, Twitter, Insta etc and these days Machines are generating data which overall creating a Huge Volume of data which cannot be handled easily through traditional databases, so after completion of this course you'll be able to do that using HADOOP as your platform.
  • Able to crack Cloudera CCA 175 Certification

Course content

1 section27 lectures28h 55m total length
  • Big Data and Hadoop Introduction1:03:24
  • Hadoop Framework57:11

    Explore big data concepts and the Hadoop ecosystem, focusing on volume, velocity, and value, and see how distributed, parallel processing handles large data using real-world examples.

  • Hadoop Ecosystem1:08:53

    Explore the Hadoop ecosystem and its master-slave architecture, learn how data is split into 128-byte blocks with a replication factor for reliability, and managed within a scalable framework.

  • HDFS1:06:29

    Explore the hdfs architecture within the big data and Hadoop ecosystem, detailing replication factors, block distribution, backups, fsimage, heartbeats, and how clusters recover from failures.

  • Magic Boxes, Sqoop and Flume1:01:27

    Explore the big data and Hadoop ecosystem, focusing on storage, processing, and distributed architectures, with practical insights into data movement using Sqoop and Flume on a Hadoop cluster.

  • NameNode, DataNode and JournalNode1:03:29

    Explore the Hadoop big data ecosystem with NameNode, DataNode, and JournalNode, learn how MapReduce handles large data, replication, and scalability, and compare SQL-based and non-SQL data processing.

  • Input output operations, Ram and HDD, pros and cons57:54

    Explore input-output operations in big data with ram and hdd considerations, outlining pros and cons of storage strategies, Hadoop's hdfs architecture with name node, data nodes, and metadata backups.

  • MapReduce Theory 1.158:22

    Explore MapReduce theory in big data and Hadoop, focusing on replication, active and backup name nodes, heartbeat monitoring, and fault-tolerant data across a cluster.

  • MapReduce Theory 1.21:01:55

    Explains how RAM and permanent storage interact, why input-output operations slow processing, and how 24-hour backups protect data in a big data cluster.

  • MapReduce Theory 1.357:57

    Explore MapReduce theory and big data concepts through practical examples, focusing on map and reduce phases, input-output models, and performance trade-offs.

  • Combiner Approach with pros and cons1:04:49

    Discover how the combiner in mapreduce reduces intermediate data, and analyze its pros and cons while examining input/output bottlenecks, RAM versus disk storage, and network transfer impacts.

  • Practical : Sqoop with MySQL58:29
  • Visit to Cloudera Machine59:39

    Explore how to access the Cloudera machine, set up Oracle VM VirtualBox, and practice MySQL commands like show databases, use, create database, and show tables, noting RAM and 64-bit requirements.

  • Sqoop commands with introduction to Linux commands as well59:58

    Explore Sqoop commands and basic linux commands to transfer data from MySQL to Hadoop using Sqoop, connecting via Cloudera, with hands-on guidance on listing databases and importing tables.

  • Sqoop commands1:00:44

    Master Sqoop commands to import data from an RDBMS like MySQL into Hadoop via HDFS, and export data back to an RDBMS, handling primary keys and blank-table setup.

  • Basics of core Java, introduction to eclipse, MapReduce Coding56:46

    Explore core Java basics, introduce Eclipse setup, and dive into MapReduce coding concepts like mappers, reducers, key-value pairs, and word counting.

  • MapReduce Coding only1:57:59

    Set up Eclipse with the Java Development Kit and Cloudera Hadoop jars, import the libraries, and build a MapReduce program featuring mapper, reducer, and a word count example.

  • Hive theory1:03:54
  • Word Count Processing time Comparison between MapReduce and Apache Spark7:51

    Compare word count performance between MapReduce and Apache Spark using a 90-byte file in a Hadoop environment, noting MapReduce takes a minute while Apache Spark finishes in seven to eight seconds.

  • Hive: connecting, loading, defining delimiters1:05:09

    Connect to Hive, create a database and table, and load data from a file into Hive table using delimiters. Explore the delimiter ctrl-a and how Hive and Impala handle loading.

  • Hive: coding1:00:13

    Master Hive coding for big data by learning how to create delimited tables, define schemas, and load data in a Cloudera Hadoop environment, with comparisons of Derby, MySQL, and Impala.

  • Hive to Impala and Beeline1:07:22

    Compare hive and impala, highlighting impala's speed advantage and mapreduce avoidance for SQL queries in a Hadoop environment. Learn to use Beeline to run these queries efficiently.

  • Comparing Performance Time between Hive and Impala6:56
  • Hive : partitioning and bucketing part 11:06:31
  • Hive : partitioning and bucketing part 2 ( 2 hours lecture)2:14:00

    Explore Hive partitioning and bucketing with practical examples on Amazon data, showing how partitioning and bucketing improve big data queries and performance.

  • YARN, HBASE, OOZIE1:19:23
  • FINAL PROJECT ON REAL DATA SET1:29:01

    Hello guys !

    Great work reaching up to here, for the project thing, some '.mod' are required as you can see in the lecture but those files are not supported here, so I would request you to please send me an email on ' sahebsinghchaddha@gmail.com ' if you need all the resources for the Project, I;ll main you personally.

    Thanks!

Requirements

  • Well, it's say that you need to have a knowledge of basics of Core Java and SQL, but when taking course from me, you don't require anything but just English to understand my lectures because everything will be covered here.
  • Get familiarize with Cloudera CDH.
  • Work on Hadoop tools like, Hive, MapReduce, Sqoop, Impala, Beeline, etc
  • NO JAVA REQUIRED IN MY COURSE
  • Introduction to Apache Spark and Scala, Machine Learning Libraries ( MLLib ), etc

Description

This is an interactive lecture of one of my Big data and Hadoop class where everything is covered from the scratch and also you will see students asking doubts so you can clear those concepts here as well.

Students will be Able to crack Cloudera CCA 175 Certification after successful completion and with little practice.

Tools covered :

1. Sqoop

2. Flume

3. MapReduce

4. Hive

5. Impala

6. Beeline

7. Apache Pig

8. HBase

9. OOZIE

10. Project on a real data set.

Who this course is for:

  • Students who want to step into Big Data, want to know how to Analyse, work on and manage it.