
Learn to write and run Apache Spark code using the Databricks online environment, using the DataFrame API and SQL for scalable big data processing.
Create a Databricks community account to access a free single-node cluster for learning Apache Spark. Sign up at Try Databricks and explore notebooks and clusters.
Upload the dataset into the Databricks file system via the data pane, unzip the Retailer package, and place data under file store/tables/retailer and images under retailer/images.
Import a Databricks archive of notebooks into your workspace, unzip the DBC file, and import it to follow along with course notebooks in Databricks.
Navigate the Databricks workspace user interface, create a cluster with a Databricks runtime for Apache Spark, and attach notebooks to run Spark code for data cleansing and machine learning.
Apache Spark runs on a cluster managed by a cluster manager and node managers, with a Spark driver requesting executors to run tasks across the cluster.
Import a Databricks notebook and run Spark code to build a customer data frame; find customers born on the same day, month, and year, and count those sharing your year.
Create a Spark cluster and notebook, then load customer.csv into a dataframe using Spark session with read.csv or read.format('CSV').load, and display the dataframe in Databricks.
Learn to select specific columns from a data frame using the select method, with string names or column objects, and display results using show in Databricks.
Understand how a dataframe's schema defines column names, data types, and nullability, view it with print schema, and explore the struct type of fields with name, data type, and nullability.
Learn to specify a data frame schema with a ddl-formatted string and manually define an address data frame using spark session, header option, and a vertical bar delimiter.
Define a data frame schema using a ddl-formatted string for the customer address data, then apply it with the data frame reader and print the schema in ddl format.
Explore how a dataframe, Spark's core data structure, organizes data in rows and columns with a schema, enabling distributed processing across executors via partitions in memory.
Rename dataframe columns to clearer names using the with column rename method, returning a new dataframe with updated columns while the original remains unchanged.
Filter out invalid birth records from a customer data frame using Spark's filter and where methods, handling nulls, valid ranges, and equality checks to clean the data set.
Join two dataframes on a join expression to combine left and right dataframes, such as customers and addresses, using inner join when keys are missing or null in Spark.
Learn to perform an inner join between dataframes in Spark by defining a join expression, joining left and right datasets, and selecting useful columns for a readable result.
Learn to count rows in a data frame using the data frame count method or spark sql functions count, including star counts and column-based counts with null handling.
Learn to compute the minimum wholesale cost from an item data frame, rename the result, and use an inner join to retrieve all products at the minimum wholesale cost.
Compute the average quantity with spark: use average or mean, or sum divided by count on the store sales data frame; alias results and view min, max, and mean values.
Group data by city (and salutation) with the group by method to create groups, then count customers in each group, using Apache Spark.
Group by customer ID in Apache Spark and perform aggregation to compute distinct item counts, total quantities, and total spend in the store sales dataframe.
Compute the total extended sales price per item for manufacturer 128 in November by joining date, store sales, and item data, then group by year, month, brand, and item.
Explore how Apache Spark executes data with dataframes, partitions, and executors, using lazy transformations, narrow and wide dependencies, and shuffles, with actions triggering jobs.
Create a temporary view from a data frame and run Spark SQL queries to retrieve results, illustrating equivalence with the data frame API and showcasing view-based querying.
Register a data frame as a global temporary view to share across notebooks, using global_temp prefix and create or replace global temp view, with behavior across sessions and cluster detach.
Create and manage tables with Spark SQL, distinguishing external unmanaged tables from managed ones, and understand how metadata and data are stored and affected by drops.
Learn how the select clause returns columns and how select expressions combine values and functions in Spark. See how the select expression builds dataframes from SQL.
Explore filtering customers born in January in Canada, using where with equality and inequality, and normalize strings with lower and trim to handle case and spaces.
Learn SQL null handling by using is null and is not null predicates to filter rows, avoiding equality checks with null, and applying these concepts to real tables.
Explore SQL aggregations with count, sum, avg, mean, min, and max, including null handling and grouping concepts using a customer table and store sales data.
Use the SQL group by clause to create groups by store key and item key. Apply sum to compute total net profit per group and handle null values.
Learn how to filter groups with the having clause alongside group by, using aggregation like count(*) to count customers by birth year and country, with examples and null handling.
Explain how the right outer join returns all rows from the right table and fills left nulls. Highlight data cleaning for invalid customer IDs before reporting.
Learn how to use the case expression to categorize store sales price data with conditional logic, returning below average, above average, or unknown based on min, average, and max prices.
Describe building a Databricks notebook to report total web sales for selected items and zip regions in 2001 Q2, using joins, filters, and a bar chart.
Learn to convert literal values to Spark types with the lit function, then add and rename a new column in a data frame and inspect its schema.
Learn to work with boolean expressions in Apache Spark to filter dataframes using where or filter. Create boolean columns and combine conditions with and and or to refine results.
Explore strings in spark using data frames and SQL, performing contains, length, trim, lower and upper, init cap, and padding, plus filtering, counting, and distinct queries on the item table.
learn to work with datetimes in apache spark by converting strings to dates with to_date, using lead, and computing date differences with date_diff and month_between.
Create struct columns in Apache Spark. Use struct and select expressions to build nested fields, alias them, access with dot notation or get_field, then run sql or dataframe queries.
Explore arrays in Apache Spark by building a data frame of customers and their item lists, and use collect_set, explode, and array_contains for analysis.
Create and query a map column in Apache Spark with category as key and product name as value, using trim, then explode and rename to category and product name.
Learn to handle null values in Apache Spark using data frame functions to drop rows with any or all nulls, fill or replace missing values.
Apply Apache Spark fill to replace nulls with defaults by column type, use maps for per-column values, and use replace to convert yes/no to 1/0 in the preferred customer flag.
Learn to read CSV files with the DataFrameReader in Spark, configuring format, schema, and header options. Manage malformed records with modes such as permissive, fail fast, and drop malformed.
Learn to read JSON files with Apache Spark using the Spark session data frame reader, handling single-line and multi-line JSON, and applying explicit schema, date format, and nested field queries.
Use the dataframe writer to save a dataframe to csv or json with options like header, separator, and compression, and partition by year and month.
learn how to create a dataframe manually
Welcome to this course on Databricks and Apache Spark 2.4 and 3.0.0
Apache Spark is a Big Data Processing Framework that runs at scale.
In this course, we will learn how to write Spark Applications using Scala and SQL.
Databricks is a company founded by the creator of Apache Spark.
Databricks offers a managed and optimized version of Apache Spark that runs in the cloud.
The main focus of this course is to teach you how to use the DataFrame API & SQL to accomplish tasks such as:
Write and run Apache Spark code using Databricks
Read and Write Data from the Databricks File System - DBFS
Explain how Apache Spark runs on a cluster with multiple Nodes
Use the DataFrame API and SQL to perform data manipulation tasks such as
Selecting, renaming and manipulating columns
Filtering, dropping and aggregating rows
Joining DataFrames
Create UDFs and use them with DataFrame API or Spark SQL
Writing DataFrames to external storage systems
List and explain the element of Apache Spark execution hierarchy such as
Jobs
Stages
Tasks