
Discover the Spark Scala coding framework and testing, with hands-on unit testing using JUnit and Scala test, building a production ready data pipeline that interfaces with Hadoop, Hive, and PostgreSQL.
Discover distributed storage and computing with Spark, and compare it to MapReduce. Learn to build Spark Scala applications using DataFrames, Spark SQL, and Hive on HDFS and cloud storage.
Set up a Windows environment for Spark, Scala, and Hive with IntelliJ. Learn to install JDK 11, IntelliJ, winutils, and configure Spark core and SQL dependencies for a Maven project.
Master IntelliJ setups on Windows and Mac. Learn that older Java and Scala code runs on newer versions without issues, keeping concepts and code consistent.
Master the basics of Scala by creating a Maven project, configuring Scala, and learning var versus val, type inference, arrays, lists, maps, tuples, and loops.
Connect a Spark application to PostgreSQL via JDBC, configure user credentials and URL, and fetch data from the future schema's Future X Course Catalog table into a DataFrame.
Import a spark scala project into IntelliJ from existing sources, configure maven with pom.xml, set JDK 1.8 and scala 2.11.8, and resolve dependencies.
Organize a Spark Scala project by creating a common package and objects to initialize the Spark session and centralize PostgreSQL properties, enabling methods to fetch data frames from PostgreSQL tables.
Configure log4j2 and SLF4J logging in a Maven-based Scala app with a Log4j2 properties file. Replace println with logger.info and experiment with info, warn, and error levels for structured logs.
Read data from a Hive table with Spark SQL, apply transformations to replace nulls with unknown or zero, and write the results to a PostgreSQL table.
Learn how to read json configuration with typesafe config, fetch the Postgres target table from a json file, and control data transformations via json-driven settings.
Learn how to pass and read command line arguments in a Spark program, debug with breakpoints in IntelliJ, and control program flow based on environment like dev or production.
Explore unit testing in scala using spark transformer tests, catching null pointer exceptions, asserting results, and leveraging fixtures and matchers to validate data frame transformations.
Throw and catch custom exceptions in Scala, intercept error messages, and validate data with tests that trigger a null course name in a Spark Scala workflow.
PySpark course link with coupon
https://www.udemy.com/course/pyspark-python-spark-hadoop-coding-framework-testing/?referralCode=F624A76C9626B8884BE3
Thank you for enrolling in the Spark Scala course; rate it if you found it valuable, and explore resources, exclusive coupons, YouTube content, and educational blogs for further learning.
Explore how Hadoop enables distributed storage and computing through hdfs and mapreduce, with data nodes and a name node coordinating a scalable cluster.
Explore how HDFS distributes large data sets across a cluster using 128 MB blocks and a replication factor of three for fault tolerance, with trade-offs between storage and resilience.
Sign up for a Google Cloud free tier account and explore data engineering and machine learning with $300 in free credits, then monitor usage in the billing console.
Create a production-grade Hadoop and Spark cluster on Google Cloud Dataproc by configuring a three-node setup, enabling the Dataproc API, and SSH into the master node.
Learn to store and manage files in HDFS using Hadoop fs commands on a Dataproc cluster, including creating directories, putting files, listing, getting, and removing files.
Discover Hive, a SQL-like tool that processes data in HDFS by translating queries to MapReduce. Learn how Hive stores metadata in metastore and uses tables and Hiveql to query data.
Query HDFS data with Hive on a Dataproc cluster. Map CSV data to Hive tables, both internal and external, and run MapReduce jobs via Yarn.
This course bridges the gap between academic learning and real-world application, preparing you for an entry-level Big Data Spark Scala Developer role. You'll gain hands-on experience with industry best practices, essential tools, and frameworks used in Spark development.
What You’ll Learn:
Spark Scala Coding Best Practices – Write clean, efficient, and maintainable code
Logging – Implement logging using Log4j and SLF4J for debugging and monitoring
Exception Handling – Learn best practices to handle errors and ensure application stability
Configuration Management – Use Typesafe Config for managing application settings
Development Setup – Work with IntelliJ and Maven for efficient Spark development
Local Hadoop Hive Environment – Simulate a real-world big data setup on your machine
PostgreSQL Integration – Read and write data to a PostgreSQL database using Spark
Unit Testing – Test Spark Scala applications using JUnit, ScalaTest, FlatSpec & Assertions
Building Data Pipelines – Integrate Hadoop, Spark, and PostgreSQL for end-to-end workflows
Bonus – Set up Cloudera QuickStart VM on Google Cloud Platform (GCP) for hands-on practice
Prerequisites:
Basic programming knowledge
Familiarity with databases
Introductory knowledge of Big Data & Spark
This course provides practical, hands-on training to help you build and deploy real-world Spark Scala applications. By the end of this course, you’ll have the confidence and skills to build, test, and deploy Spark Scala applications in a real-world big data environment.
This course uses high-quality AI-generated text-to-speech narration to complement the powerful visuals and enhance your learning experience.