
Learn big data testing with Spark, Scala and MongoDB through streaming topics and practical data examples demonstrated in this session.
Explore Spark components, the open-source cluster computing engine that powers fast batch and streaming processing, with Spark's Scala and Java APIs and MLlib for machine learning.
Trace the history of Spark from its 2009 origins to its open source foundation. Learn how Spark enables a flexible, accessible big data pipeline.
Explore different ways to define variables in Scala, including declaring and assigning values, and learn the rules that govern how and when a value can be set.
learn how to set up a Cloudera environment for big data work by installing a 64-bit virtual box, downloading Cloudera open source software, and configuring the environment.
Participate in a practical session on using variables in Scala, including defining and assigning values and reading back results, within the context of big data testing with Spark and MongoDB.
Learn to run basic Spark commands in the Spark shell with Scala, observing how the shell executes commands and displays results.
Load data from a file in Scala through a practical session, read the source content, and produce line-by-line output.
Explore practical techniques for working with lists in Scala, including creation and basic manipulation to manage data in Scala.
Build your first Scala program in Eclipse by using the Scala IDE, typing Scala code, and understanding basic Scala workflows integrated with Java.
Analyze practical variable implementation in Scala during the big data testing course, focusing on memory, value reassignment, and exception in Spark, Scala, and MongoDB workflows.
Engage in a practical session on variable implementation to solidify foundational skills. Apply concepts and execute steps to strengthen big data testing workflows with Spark, Scala, and MongoDB.
Engage in a practical session on creating and applying functions in Scala, exploring function design, evaluation, and real-world usage for big data testing.
Explore spark architecture by examining cluster managers, the driver, and executors, and how applications run across multiple machines with URLs and incoming connections.
Learn how rdds in spark are immutable, support lazy evaluation, and are fault-tolerant, with partitioning and caching that enable memory-efficient, parallel big data processing.
Learn the most important acronyms and keywords used in Spark programs, including Spark and Scala concepts, cluster and resource manager concepts, Yarn, spot on, and Parquet data formats.
Learn how to work with collections in Scala, including building and reading a list of six values, using count and map, and configuring dependencies for Spark in a project.
Learn how to work with file data using Spark context by creating a dataset from a text file, applying map and collect, and running a Scala Spark application.
Learn how to read and work with JSON file format data using Spark and Scala, with practical steps to parse, structure, and extract values from JSON files.
Learn to read and work with parquet file format data, loading datasets, displaying them in a table, and addressing formatting and labeling challenges.
Learn how to read data from Hive using HiveContext in Spark, including connecting to Hive, selecting tables, and loading data for processing.
Explore how to apply conditions and filters with HiveContext in Spark, enabling precise data selection for big data testing with Spark, Scala, and MongoDB.
Learn what MongoDB is, its open-source, cross-platform document database nature, its scalability and indexing, and the CRUD operations across collections and documents.
Install MongoDB on an Ubuntu machine by updating the package list, installing the MongoDB package, and starting the mongod service to run locally.
Learn how to create a MongoDB database, add collections and documents, and insert values to store and retrieve data.
Learn to insert documents into a MongoDB collection using a defined schema, with fields like name and age, and to query the data using find and find one.
Explore how to read data from MongoDB collections using dot notation, apply limits (including limit 1), and filter with comparison operators like $lt and $gt to retrieve relevant records.
Learn how to update data in a MongoDB collection by matching documents, changing field values, and verifying updates with queries. Explore updating single or multiple records and viewing results.
Learn how to delete data in MongoDB by removing documents from a collection, using delete and remove operations, and understanding the effects on your dataset.
Explore big data testing fundamentals, including the three main concepts: volume, velocity, and variety, plus variability and complexity, with interview-ready insights and examples of technologies and testing tools.
Discover how high provides a SQL-like interface on the Hortonworks data platform, translating queries into MapReduce via Hadoop, with UDFs, UDTFs, and HQL for data summarization and analysis.
This course is for Testing profile candidate who wanted to build there career into Big Data Testing. So I have designed this course so they can start working with Spark, Scala and MongoDB into big data testing. All the users who are working in QA profile and wanted to move into big data testing domain should take this course and go through the complete tutorials.
I have included the material which is needed for big data testing profile and it has all the necessary contents which is required for learning Spark, Scala and MongoDB.
It will give the detailed information for different Spark, Scala and MongoDB commands which is needed by the tester to move into bigger umbrella i.e. Big Data Testing.
This course is well structured with all elements of different Spark, Scala and MongoDB commands in practical manner separated by different topics. Students should take this course who wanted to learn Spark, Scala and MongoDB from scratch.