
Explore hands-on preparation for the Cloudera SCCA 175 exam using Apache Spark to solve big data problems, with a ready-to-use VM, datasets on HDFS, and Zeppelin notebooks.
Install Oracle VirtualBox, download and import the Verulam blue VM, and boot a Spark Hadoop stack with Hive and Zeppelin to run on your PC.
Boot the virgin blue VM via the Oracle VM shortcut and start the Farallon Blue cca 175 environment. Resize the desktop if needed, then use Spock, Hive, and Zeppelin notebook.
Spin up the Hadoop cluster on the blue pvm, then verify seven daemons, including namenode, datanode, resource manager, node manager, and hive metastore server, are running.
Open spark-shell and the 30 problems to practice exam-style data analysis by using the desktop problems icons to view the problem set.
Launch the Apache Zeppelin notebook from the Zeppelin icon, access and write solutions, run and rerun notebooks, follow a Zeppelin tutorial, and use Spark SQL queries on data.
Transform data from a park file format into a tab-delimited text file, zip-compress the output, and include only cardholder name, issuing bank, and issue date in the HDFC directory.
Generate a report of patients registered in baltimore gp practices and store the results in a metastable table named q14 in the gp database with snappy compression.
Transform csv data in hdfc online retail directory to adjacent file format, filter unit price above ten, prefix with a dollar sign, and save as JOOSEP-compressed in the hfs directory.
*Important Notice*
This course has been retired and is no longer receiving support. Originally designed to help students pass the now-retired Cloudera Certification exams, the material remains useful for those wanting to practice their skills on Spark and Hadoop clusters. However, its primary focus was certification preparation, which many students successfully completed.
Prepare for the transform, stage and store section of the CCA Spark & Hadoop Developer certification and helps pass the CCA175 exam.
Students enrolling on this course can be 100% confident that after working on the problems contained here they will be in a great position to pass the transform, stage and store section of the CCA175 exam.
As the number of vacancies for big data, machine learning & data science roles continue to grow, so too will the demand for qualified individuals to fill those roles.
It’s often the case the case that to stand out from the crowd, it’s necessary to get certified.
This exam preparation series has been designed to help YOU pass the Cloudera certification CCA175, this is a hands-on, practical exam where the primary focus is on using Apache Spark to solve Big Data problems.
On solving the problems contained here you’ll have all the necessary skills & the confidence to handle any transform, stage & store related questions that come your way in the exam.
(a) There are 30 problems in this part of the exam preparation series. All of which are directly related to the transform, stage & store component of the CCA175 exam syllabus.
(b) Fully worked out solutions to all the problems.
(c) Also included is the Verulam Blue virtual machine which is an environment that has a spark Hadoop cluster already installed so that you can practice working on the problems.
• The VM contains a Spark stack which allows you to read and write data to & from the Hadoop file system as well as to store metastore tables on the Hive metastore.
• All the datasets you need for the problems are already loaded onto HDFS, so you don’t have to do any extra work.
• The VM also has Apache Zeppelin installed with fully executed Zeppelin notebooks that contain solutions to the problems.