Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
An Advanced Guide for Apache Hive: A Hadoop Ecosystem Tool
Rating: 3.6 out of 5(35 ratings)
193 students

An Advanced Guide for Apache Hive: A Hadoop Ecosystem Tool

Learn Apache Hive SQL Layer on Apache Hadoop
Last updated 7/2017
English
English [Auto],

What you'll learn

  • Create Databases, Table
  • Create Hive Datawarehouse
  • Installing, managing and monitoring Hadoop cluster on cloud
  • Writing UDFs to solve the complex problems
  • Querying and managing large datasets that reside in distributed storage
  • Transforming unstructured and semi-structured data into usable schema-based data
  • Writing HiveQL statements for the same as you write MapReduce program in any host language

Course content

10 sections51 lectures7h 13m total length
  • 1.1 What is Apache Hive1:19

    Apache Hive enables querying and managing large datasets in distributed storage, offering a SQL-like dialect, formats such as text, RCFile, ORC, Parquet, compression with Snappy, and built-in functions.

  • 1.2 Prerequisites for this Course0:35

    Identify the prerequisites for learning Apache Hive, including knowledge of distributed applications, familiarity with Linux systems, and prior experience with relational databases.

  • 1.3 Syllabus of this Course2:30

    Explore the Apache Hive syllabus in the Hadoop ecosystem, covering Hive architecture, DML and DQL, views and joins, data formats from text to RCFile and JSON, plus partitioning and bucketing.

  • 1.4 Comparison of Hive with HBase and PIG1:49

    Compare Hive with HBase and Pig to understand data storage and processing in the Hadoop ecosystem; explore structured vs semi-structured data, partitions, and reporting capabilities.

Requirements

  • You should have basic knowledge of Big Data
  • You should have basic knowledge of Hadoop
  • You should have basic knowledge of RDBMS
  • You should have basic knowledge of SQL
  • You should have basic knowledge of MapReduce

Description

Hive is a SQL Layer on Hadoop, data warehouse infrastructure tool to process structured data in Hadoop. This course on Apache Hive includes the following topics:

  • Using Apache Hive to build tables and databases to analyse Big Data
  • Installing, managing and monitoring Hadoop cluster on cloud
  • Writing UDFs to solve the complex problems
  • Querying and managing large datasets that reside in distributed storage
  • Transforming unstructured and semi-structured data into usable schema-based data
  • Writing HiveQL statements for the same as you write MapReduce program in any host language
  • Solving real case studies and work on Projects with live data from Twitter

Who this course is for:

  • Any student who has hunger to learn
  • Any professional or student who want to make career in the field of Big Data and Hadoop