
Explore how Apache Hadoop powers big data processing, learn its framework and distributions, and master installation, configuration, and production cluster administration.
Understand big data by defining it, exploring its features like unstructured data and large volume, and learn how Hadoop and file systems enable scalable storage, processing, and analytics.
Explore the fundamentals of a Hadoop cluster, including the Hadoop distributed file system, name node, data nodes, and resource manager, plus how replication, blocks, and containers enable scalable, fault-tolerant analytics.
MapReduce enables parallel processing by mapping input data to key-value pairs, shuffling and reducing to produce word counts, log file analysis, and page rank on large data sets.
Explore the basics of cluster administration in Hadoop, covering Yarn and Mesos cluster managers, the resource manager and node manager, application master, and container workflows.
Explore Hadoop components across core libraries, data integration, data access and transformation, and management and monitoring, including HDFS, YARN, MapReduce, scoup, and Hive.
Discover how bash and the terminal empower Hadoop administrators to manage clusters, check node status, start or stop services, and automate tasks via command line interfaces like hive and scoop.
Learn to install Hadoop on Linux, configure user properties and OS support, choose installation modes (local, pseudo, distributed), set up passwordless SSH, and automate with a bash script.
Set up passwordless ssh across cluster nodes by generating public/private keys, distributing the public keys to authorized_keys on each machine, and securing ssh permissions.
Install Hortonworks HDP 2.4 offline by hosting depositories as tarballs on a local web server, ensure passwordless SSH between hosts, and prepare repositories for all nodes.
Plan capacity by estimating data volume, retention policy, and workload type to size a homogeneous Hadoop cluster with threefold replication, and install Apache Ambari on a master host.
Install Ambari Server and configure a Hadoop HDP cluster using the installation wizard, register hosts, select core services, and tune memory and DFS settings for reliability.
Explore the Apache Hadoop cluster management console to provision services, manage hosts and dashboards, monitor metrics, and perform maintenance, decommissioning, and permission and group management.
Explore web console operations for managing an Apache Hadoop cluster: add and remove hosts, provision and decommission services, and implement rack awareness for fault tolerance.
Explore the resource manager user interface in Hadoop, learn how the capacity scheduler uses hierarchical subqueues to share cluster resources, and review application states and logs in YARN.
Create a user home directory in HDFS, assign ownership, and set permissions with chmod and chown using the Linux-like DFS CLI; test access by copying files to and from HDFS.
Learn how HDFS snapshots provide point-in-time data protection and how access control lists enforce granular permissions with automatic propagation and restoration workflows.
Enable Hadoop high availability by configuring an active and standby name node with automatic failover and synchronized edit logs with shared storage, ensuring 3 zookeeper servers not in maintenance mode.
Hortonworks Data Platform Certified Administrator, is a Hadoop System Administrator who is capable and responsible for installing, configuring, and supporting an HDP Cluster. This is a hands-on performance-based exam which requires some competency and Big Data expertise. Tasks are advocated on AWS instances which must be completed by the candidate within a limited time frame.