
Gain an entry-level overview of big data and Apache Spark, including Spark's resilient distributed datasets and the programming language used to create Spark.
This lecture discusses:
This lecture discusses:
This lecture discusses:
Explore what Apache Spark is, its history, and why you should use it in enterprise environments, as this session dives deeper into Spark's core concepts.
This lecture discusses:
This lecture discusses:
This lecture discusses:
Conclude section two by tracing the history and creation of Apache Spark and explaining why to use it, while detailing its features driving adoption and previewing deployment and infrastructure concepts.
Explore deployment modes for Apache Spark and learn how to install standalone Spark on local machines. Review the Spark framework's main concepts and Spark applications.
This lecture discusses:
This hands-on exercise will guide you through:
This lecture discusses:
This lecture discusses:
This lecture discusses:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses:
This hands-on exercise gives practice with:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses:
This hands-on exercise provides practice with:
This hands-on exercise provides practice with:
This lecture discusses:
This hands-on exercise provides practice with:
Explore the resilient distributed datasets in Apache Spark, focusing on the RTD framework, fault tolerance, lazy evaluation, and persistence through transformations. Learn about actions, creating RDDs, and shared variables.
This lecture discusses:
This lecture discusses:
This hand-on exercise provides practice with:
This lecture discusses:
This hands-on exercise provides practice with:
This lecture discusses several topics relating to RDD key/value pairs:
This hands-on exercise provides practice with:
This lecture discusses:
This lecture discusses:
This hands-on exercise provides practice with:
Conclude section 5 by highlighting Spark's resilient distributed datasets, transformations and actions, plus fault tolerance, lazy evaluation, persistence, and shared variables.
What is Apache Spark?
Apache Spark is the next generation open source Big Data processing engine. Spark is designed to provide fast processing of large datasets and high performance for a wide range of applications. Spark enables in-memory cluster computing which greatly improves the speed of iterative algorithms and interactive data mining tasks.
Course Outcomes
'Introduction to Apache Spark' includes illuminating video lectures, practical hands-on Scala and Spark exercises, a guide to local installation of Spark, and quizzes. In this course, we guide students through:
Upon completion of the course, students will be able to explain core concepts relating to Spark, understand the fundamentals of coding in Scala, and execute basic programming and data manipulation in Spark. This course will take approximately 8 hours to complete.
Recommended Experience
Programming Languages recommended for this course:
Recommended for:
For students unfamiliar with Big Data and Hadoop, the course will provide a brief overview of each topic.
Why Adastra Academy?
Adastra Academy is a leading source of training and development for Information Management professionals and individuals interested in Data Management and Analytics technology. Our dedication to identifying and mastering emerging technologies guarantees our students are the first to have access to these quality courses. For an exceptional learning experience, our programs include hands-on labs and real world examples allowing students to easily apply their new knowledge.