
Explore tidymodels in R for regression, blending theory with practical coding tasks. Prepare for a two-part series that expands topics in statistical modeling and machine learning.
Learn the course structure, including theoretical lectures with slides and practical coding sessions. Understand how assignments, exercises, PDFs, and datasets are used throughout the course.
exercises and assignments test your coding skills with provided code and walkthroughs; you can work solo or use the solution script, and later sections add more assignments.
Code along with the provided R scripts or review them after the video, and access the GitHub repository with downloadable code and section YAML details.
Build a foundation in data science by linking statistics and machine learning, covering regression, classification, and clustering, and start with R, RStudio, and the tidyverse for tidymodels.
Data science uses data to answer questions by combining mathematics, statistics, computer science, and domain expertise to extract insights and guide predictive modeling across industries.
Understand statistics as the discipline of collecting, organizing, and analyzing data, with descriptive and inferential branches, and see how statistical learning blends statistics and machine learning to predict outcomes.
Explore how machine learning develops algorithms and models that enable computers to learn from data, identify patterns, and predict unseen future outcomes, using supervised, unsupervised, and reinforcement learning.
Explore how statistics and machine learning connect and differ, from origins in mathematics and artificial intelligence to the role of statistical modeling in evaluating predictive models.
Explore regression, classification, and clustering within supervised and unsupervised learning using tidymodels in R; learn how models predict continuous outputs, categorize data, and reveal hidden patterns.
Introduce the R programming language as the primary tool for data analysis, modeling, and machine learning, and explain RStudio as its integrated development environment for coding, debugging, and tidyverse basics.
Install and configure R and RStudio, with optional R Tools, using the course links to download the latest versions for Windows, Linux, or macOS.
This lecture introduces tidyverse and tidymodels, emphasizing tidy data and data wrangling as foundations for building predictive models in R.
Define data science, statistics, and machine learning, and show how they relate as sides of the same coin. Preview regression modeling with tidyverse and tidymodels, emphasizing data wrangling and preparation.
Explore the tidyverse basics, including dplyr and tidier, learn ggplot2 for exploratory data analysis, and practice with exercises and assignments using map functions.
Explore the tidyverse and tidymodels ecosystem, transforming raw data into tidy data for efficient exploration and modeling with dplyr, ggplot2, readr, and purrr.
Learn to wrangle data with dplyr from tidyverse, using verbs like select, mutate, filter, arrange, distinct, and the forward pipe to chain operations for tidy data frames.
Learn dplyr basics in tidyverse, including selecting, renaming, mutating, filtering, arranging, and deduplicating data, using the mpg dataset and pipe chaining in R.
Explore dplyr summarization in R tidymodels part 1 by learning summarize, group by, and aggregation functions (min, max, mean, sd, var, n) for grouped data.
Explore dplyr summarise and group_by in the tidyverse with mpg and h flights datasets, computing averages, counts, distinct manufacturers, and percentage of canceled flights per carrier, then sort results.
Explore tidyr basics with pivot longer and pivot wider to convert between wide and long data formats, and learn how to reshape tables and nest data for tidy modeling.
Explore tidyr fundamentals by converting a long car data frame to wide with pivot wider, then returning to long with pivot longer, including filtering and max city consumption.
Practice data manipulation with dplyr and tidyr to compute total distance flown and total flights cancelled per carrier, then analyze weekly flight patterns using tidyverse tools.
Explore practical dplyr and tidyr solutions with step-by-step examples using cars and flights data, including counting, filtering, grouping, summarizing, and pivoting with lubridate and stringr.
Explore exploratory data analysis and ggplot2 within the tidyverse, learning to visualize and clean data, build insight, and create layer by layer plots using the grammar of graphics.
Master ggplot2 by learning core plots for exploratory data analysis, including histograms, density plots, bar charts, scatter plots, and box plots, with practical mapping and layering explained.
Learn to create histograms, density plots, bar plots, and scatter plots with ggplot2, customize axes and titles, and add a regression line using cars data set and flight data set.
Demonstrates ggplot2 techniques to build and customize statistics plots, including bar, scatter, box, violin, and line plots with color fills, legends, and viridis scales using car data and flight trends.
Learn to use ggplot2 for exploratory data analysis with dplyr, creating histograms, bar plots, scatter plots, facets, and heat maps from cars and flight data.
Watch a solution walkthrough of ggplot2 eda, loading data and building histograms, barplots, and scatter plots. Adjust bin width, colors, themes, and axis labels across car and flight examples.
Learn functional programming in R with purrr map functions, applying a function over lists or vectors and replacing base R apply, while exploring iteration without loops.
Explore purrr in R by applying map over data frame columns, computing means and maxima for numeric columns, and nesting data frames by manufacturer for modeling.
Explain tidyverse tools, including dplyr and tidier, and demonstrate ggplot2 for exploratory data analysis, pivoting data between long and wide formats, and map-based functional programming, preparing for Tidymodels foundations.
Explore two data sets: insurance costs and bike sharing, examining age, BMI, sex, smoker, region, and weather, and practice loading, summarizing, and visualizing in R with tidymodels.
Carry out the assignment walkthrough by loading insurance and bike rental data, summarizing age, bmi, smoker, region and costs, visualizing distributions, and emphasizing exploratory data analysis before tidymodels modeling.
Explore Tidymodels after practicing classical R linear regression with lm, outline the machine learning workflow, and guide you to build your first Tidymodels model.
Explain linear regression with a height–weight example, estimating intercept and slope by least squares and examining residuals and noise. Implement the model in R with lm, and introduce tidymodels.
Learn to perform a simple linear regression in R in RStudio: import height-weight data, fit a model, predict weights, visualize the regression line, and read the model summary.
Learn to interpret the linear regression output in RStudio, including intercept and height coefficients, residuals, and metrics like p-values, r-squared, adjusted r-squared, for weight prediction.
Use faithful data to fit a linear regression predicting waiting time from eruption duration. Visualize the relationship, generate predictions and residuals, and assess normality of residuals to validate the model.
Apply linear regression to predict waiting time from eruption duration, visualize with a scatter plot and regression line, and evaluate residuals with a histogram.
Explore tidymodels from an intro perspective, learning how to define problems, prepare data, engineer features, split data, and build, tune, and evaluate regression models using parsnip, recipes, yardstick, and workflows.
Explore the tidymodels workflow, covering data splitting with train/test and cross-validation, recipe-based preprocessing, parsnip model specification, and hyperparameter tuning with tune and dials.
Learn to build tidymodels regression models using height and insurance data, split data with train-test, evaluate RMSE, and apply one-hot encoding via recipes.
Develop a tidymodels linear regression workflow with heights and weights data: load, split, create recipe and model, fit, predict on train and test, and assess RMSE.
Build three linear regression models with tidymodels on insurance data, including data loading, train-test split, and one-hot encoding. Compare models using rmse on training and testing sets.
Compare three linear regression models with tidymodels by merging RMSE results for train and test data, then visualize and interpret performance to select a simple yet effective model.
Learn tidymodels basics and the linear regression workflow, from theoretical background to hands-on modeling and statistical inference, with two datasets and an upcoming assignment.
build and evaluate linear models to predict mpg from mtcars using weight, horsepower, cylinders, and transmission; perform EDA with ggplot2 and pair plots, and split data for training and testing.
Explore the motor trend car data with tidymodels and tidyverse. Build and compare linear regression models using train-test splits, one-hot encoding, and RMSE evaluation.
Explore the diamonds dataset and apply penalized regression to predict diamond prices, covering feature engineering, feature selection, ridge, lasso, and elastic net, with hands-on coding and a house-price assignment.
Explore the diamonds data set from ggplot2 and learn how price relates to carats, cut, color, and clarity. Build predictive models starting with linear regression and penalized upgrades.
Load the diamonds dataset, create tiny, small, medium, and big sample subsets, and visualize distributions of price and features to prepare for price prediction with penalized regression.
Learn exploratory data analysis of the diamonds dataset in R tidymodels part 1, examining price against carat and other features, and applying log transformations to reveal patterns.
Explore the diamonds dataset through end-to-end eda by plotting log price against log carat and volume, using ggplot, and highlight feature engineering for price modeling with color, cut, and clarity.
Load the diamonds data set, create log price, log volume, volume, and aspect ratio through feature engineering, and build a dataprep diamonds function to prepare multiple data frames.
Explore feature engineering and exploratory data analysis in an R tidymodels solution, building new features from cut and color and examining distributions and log price and log volume.
Explore feature selection theory and practice with tidymodels in R, including filter, wrapper, embedded methods, dimensionality reduction, domain knowledge, and cross-validation for robust model performance.
Explore backward and forward feature selection with tidymodels to build and compare linear regression models, including 80/20 train-test split and data cleaning.
Explore forward and backward feature selection within tidymodels, building a full and minimal model, comparing selections with test data, and evaluating predictions using RMSE.
Explore penalized regression concepts, including ridge and lasso, to reduce overfitting and handle multicollinearity, while understanding the bias-variance trade-off and feature selection.
Explore penalized regression with tidymodels, learning how hyperparameters like lambda govern regularization. Practice cross-validation tuning for ridge, lasso, and elastic net, with normalization and workflow steps.
Explore ridge and lasso regression with tidymodels on diamonds data, using 80/20 splits, 5-fold cross-validation, normalization, and rmse-based model selection.
Explore penalized regression with tidymodels by comparing ridge and lasso, tuning the penalty on a log-scale grid, and evaluating predictions on test data while extracting coefficients.
Explore elastic net regression, blending ridge and lasso penalties in a single mixed model. Tune alpha (mixture) and lambda (penalty) in tidymodels with glmnet, balancing regularization and feature selection.
Build and compare elastic net regression on the diamonds dataset, tuning lambda and alpha, and contrast elastic net with ridge, lasso, and feature selection models.
Summarizes the diamonds data set with exploration, eda, feature engineering and selection, then tidymodels-based ridge, lasso and elastic net penalties, model comparison, and final assignment on a new data set.
Leverage tidymodels regression to predict California district housing prices from Kaggle data using six model settings, with 70/20/10 data split, five-fold CV for penalized models, and kNN imputation.
Explore a practical tidymodels workflow: import data, impute missing values with kNN, engineer features, split data into train/validate/test, and compare linear and penalized models using RMSE.
Conclude the course by reflecting on Tidymodels and machine learning workflows, outline the six-part plan, emphasize regression focus, and preview future parts that expand modeling context.
Are you ready to move beyond data wrangling and start building real predictive models in R?
This course is your next step!
Whether you're a data analyst, aspiring data scientist, or a tidyverse user seeking to enhance your modeling skills, this course introduces you to tidymodels, a powerful and consistent framework for statistical modeling and machine learning in R.
This course gives you a solid foundation in predictive modeling. We begin with the fundamentals of regression modeling and guide you through the complete modeling workflow using tidymodels:
Understand what modeling is — and how it's different from just analyzing data.
Grasp the principles of statistical learning and machine learning.
Build, validate, and interpret linear regression models.
Discover the power of penalized regression (ridge, lasso, elastic net).
Learn the concept of the bias-variance trade-off and how regularization helps.
Apply consistent, tidy workflows for:
preprocessing with recipes
modeling with parsnip
resampling with rsample
evaluation with yardstick
tuning with tune
Start using best practices like train/test splits, cross-validation, and performance metrics.
You'll walk away not just knowing how to use the tools, but understanding the modeling process itself!
Why Tidymodels?
The tidymodels ecosystem brings the same clarity, consistency, and elegance that you love in tidyverse, but for modeling.
Instead of jumping between inconsistent modeling functions and ad-hoc code, tidymodels lets you build, tune, and evaluate models using a well-structured and coherent grammar.
What You’ll Get
Clear explanations of modeling concepts
Practical coding demos using real-world data
Step-by-step modeling workflows
Downloadable R scripts and datasets
Exercises and assignments to reinforce learning
Solutions for all exercises and assignments
Lifetime access
If you’ve mastered tidyverse and now want to predict, model, and explain, then this course is your launchpad.
Enroll today and start building models the tidy way!!!