
Explore Hadoop's expanding ecosystem, from DFS to streaming and other data technologies, enabling polyglot programming, data science, and machine learning with Java, Hive, and MongoDB integrations.
Source VMs for the Projects
Explore MapReduce with Hadoop to mine existing data, using shuffle sort, NLP entity extraction, and clustering to identify duplicates and aggregate addresses for value-driven insights.
Extend the MapReduce job to output duplicates as cluster seeds and migrate data from hdfs to the local filesystem for etl analysis, then perform hierarchical clustering.
Explore how Apache Crunch and stringmetric enable efficient Hadoop clustering through MapReduce pipelines, ETL phases, and distance-based entity duplication detection.
Build a distributed streaming application with Kafka, Yarn, and Zookeeper on Hadoop 2.0, using Samza to process Twitter streams and manage Yarn containers.
Set up a Hortonworks virtual machine with Ambari to configure Storm and Kafka, build Kafka and Storm artifacts with Maven, and test the topology in local mode before cluster deployment.
Learn to run real-time stream processing in cluster mode by building Kafka and Storm artifacts with Maven profiles, configuring runOnCluster, deploying jars, and starting services with Ambari.
Learn how HDDACCESS enables semantic interoperability in healthcare apps by integrating RxNorm, SNOMED, LOINC, and ICD, and performing NCID and representation searches with Lucene, Solr, and Hadoop.
Migrate data from MySQL to Hadoop with Sqoop, use Hive as the data warehouse, and index with Solr on Lucene for fast MapReduce-accelerated queries.
Explore Hive usage for big data in healthcare, run vanilla SQL in Hive, perform joins with MapReduce, and index data in Solr using Lucene.
Learn to integrate pig scripts with ipython notebooks to run mapreduce jobs on hdfs, filter data with pandas, and analyze hadoop log data using hive and pig.
Develop and review end-to-end big data modeling workflow by loading hdfs data into pandas, exploring with dataframes, and applying logistic regression via patsy, uncovering surface temperature as a key predictor.
Explore visual analytics on big data by building a Spark-on-Yarn workflow using PySpark, Seaborn, and Spark SQL APIs to perform in-memory iterative machine learning.
Learn to set up Java dependencies for Hadoop projects, configure PySpark on a Yarn cluster, connect Spark with IPython Notebook, and prepare Python and Java environments with Gradle and Seaborn.
Explore PySpark analytics on Yarn, bridging Python to HDFS, converting Spark RDDs to DataFrames, and using Spark SQL to run SQL queries.
Demonstrates building big data analytics workflows for e-commerce using text mining: tokenization, regular expressions, sentiment analysis, named entity recognition, and cloud-based scalable processing.
The most awaited Big Data course on the planet is here. The course covers all the major big data technologies within the Hadoop ecosystem and weave them together in real life projects. So while doing the course you not only learn the nuances of the hadoop and its associated technologies but see how they solve real world problems and how they are being used by companies worldwide.
This course will help you take a quantum jump and will help you build Hadoop solutions that will solve real world problems. However we must warn you that this course is not for the faint hearted and will test your abilities and knowledge while help you build a cutting edge knowhow in the most happening technology space. The course focuses on the following topics
Add Value to Existing Data - Learn how technologies such as Mapreduce applies to Clustering problems. The project focus on removing duplicate or equivalent values from a very large data set with Mapreduce.
Hadoop
Analytics and NoSQL - Parse a twitter stream with Python, extract keyword with apache pig and map to hdfs, pull from hdfs and push to mongodb with pig, visualise data with node js . Learn all this in this cool project.
Kafka Streaming with Yarn and Zookeeper - Set up a twitter stream with Python, set up a Kafka stream with java code for producers and consumers, package and deploy java code with apache samza.
Real-Time Stream Processing with Apache Kafka and Apache Storm - This project focus on twitter streaming but uses Kafka and apache storm and you will learn to use each of them effectively.
Big Data Applications for the Healthcare Industry with Apache Sqoop and Apache Solr - Set up the relational schema for a Health Care Data dictionary used by the US Dept of Veterans Affairs, demonstrate underlying technology and conceptual framework. Demonstrate issues with certain join queries that fail on MySQL, map technology to a Hadoop/Hive stack with Scoop and HCatalog, show how this stack can perform the query successfully.
Log collection and analytics with the Hadoop Distributed File System using Apache Flume and Apache HCatalog - Use Apache Flume and Apache HCatalog to map real time log stream to hdfs and tail this file as Flume event stream. , Map data from hdfs to Python with Pig, use Python modules for analytic queries
Data Science with Hadoop Predictive Analytics - Create structured data with Mapreduce, Map data from hdfs to Python with Pig, run Python Machine Learning logistic regression, use Python modules for regression matrices and supervise training
Visual Analytics with Apache Spark on Yarn - Create structured data with Mapreduce, Map data from hdfs to Python with Spark, convert Spark dataframes and RDD’s to Python datastructures, Perform Python visualisations
Customer 360 degree view, Big Data
Analytics for e-commerce - Demonstrate use of EComerce tool ‘Datameer’ to perform many fof the analytic queries from part 6,7 and 8. Perform queries in the context of Senitment analysis and Twiteer stream.
Putting it all together Big Data with Amazon Elastic Map Reduce - Rub clustering code on AWS Mapreduce cluster. Using AWS Java sdk spin up a Dedicated task cluster with the same attributes.
So after this course you can confidently built almost any system within the Hadoop family of technologies. This course comes with complete source code and fully operational Virtual machines which will help you build the projects quickly without wasting too much time on system setup. The course also comes with English captions. So buckle up and join us on our journey into the Big Data.