
Explore tidymodels part two beyond linear regression, an upgrade from part one that introduces advanced regression algorithms and lays out course structure and materials.
Explore the global learning path for tidymodels part 2, upgrading from linear regression to more complex regression algorithms. Preview upcoming topics on classification and unsupervised learning with clustering.
Explore the course layout for tidymodels part 2, balancing theory lectures and practical coding videos in R, with downloadable slides, scripts, datasets, and assignments.
Explore the part two structure, with a single exercise per section introduced midsection to test your coding skills, and use the provided solution video or script for assignments.
Access the provided R scripts and learn to use GitHub for code versioning, including cloning and session version notes (R, RStudio, package versions) and future public access.
Explore k nearest neighbor and decision tree regression in Tidymodels, covering theory, hyperparameter tuning, and practical coding with a California housing price prediction exercise.
Introduces the k nearest neighbors algorithm as a nonlinear, non-parametric method for classification and regression, detailing k selection, distance measures, scaling, and averaging the neighbors' targets.
Learn how Tidymodels introduces the k-nearest neighbors workflow for regression, covering hyperparameter tuning of neighbor count, weighting options, preprocessing with one-hot encoding and normalization, cross-validation, and selecting the best model.
Explore hyperparameters tuning in regression and classification using grid and iterative searches within tidymodels, balancing bias and variance via cross-validation to optimize model performance.
Explore knn with tidymodels in a regression workflow using the diamonds dataset, including data prep, recipe steps, cross-validated hyperparameter tuning, and rmse evaluation.
Tidymodels part 2 demonstrates tuning a knn with a grid and cross-validation, using parallel computing to speed up training, selecting the best model by RMSE, and evaluating on test data.
Practice KNN to predict California housing prices using a full tidymodels workflow, including data prep, 80/20 split, cross-validated tuning, and RMSE evaluation.
Explore k-nearest neighbors for housing price prediction with tidymodels part 2, including data imputation, an 80/20 train-test split, normalization and one-hot encoding, and RMSE-driven hyperparameter tuning.
Explore how decision trees support regression and classification, using recursive partitioning to split data by features, minimize residual sum of squares, and prevent overfitting with stopping criteria.
Learn to optimize decision trees with pruning to improve generalization and avoid overfitting. Explore pre-pruning and post-pruning, tune max depth, min samples per leaf, and cost complexity alpha via cross-validation.
Learn to build decision trees with tidymodels using engines such as rpart, partykit, C5.0, and cubist; tune min split, tree depth, and cost complexity, and visualize with rpart.plot.
Learn to build and tune a tidymodels decision tree regression model for diamond prices, using ten-fold cross-validation and a five-level grid of min_n, depth, and cost complexity.
Explore tidymodels decision trees for regression: extract the final tree, visualize with rpart.plot, and evaluate RMSE on test data.
Explore k-nearest neighbor regression and decision trees within tidymodels, detailing the machine learning workflow, hyperparameter tuning strategies, and hands-on coding in RStudio to predict diamond and housing prices.
Train two decision tree models on the California housing prices data with tidymodels, 80/20 split, 10-fold CV, and bias optimization plus simulated annealing; compare RMSE on the test set.
The assignment walkthrough explains data preparation, an 80/20 split, and tuning a tidymodels decision tree with bias optimization and simulated annealing, evaluated by rmse on the test set.
Explore ensemble learning with random forest, bagging, and parallel computing in Tidymodels, then apply eda on the Ames data and build a random forest predictor.
Explore ensemble learning by combining multiple weak learners into a strong model to reduce variance and overfitting, using bagging and bootstrap aggregation; learn aggregation methods for regression and classification.
Explore random forest, an ensemble of decision trees that uses bootstrap samples and random feature selection to reduce overfitting, improve accuracy, and provide out-of-bag error estimates.
Learn to build a random forest in tidymodels using the ranger engine, tune key hyperparameters via cross-validation, and finalize a workflow with a complete model fit on training data.
Tune a random forest with tidymodels on the diamonds data, adjusting mtry, min_n, and trees via ten-fold cross-validation with Ranger, then finalize and evaluate RMSE.
Tune and evaluate a random forest in tidymodels, visualize tuning results, select the best hyperparameters, finalize the workflow, and assess predictions with RMSE on a price log dataset.
Explore parallel computing in tidymodels using future, doFuture, and forEach to run resamples and hyperparameter tuning across multiple cores, with plan options and memory considerations.
Explore the Ames housing data set for regression, with sale price as the target and 73 predictors. Apply eda on distributions, missing values, and correlations before a random forest model.
Explore Ames housing data with expansive exploratory data analysis, examining data structure, distributions, correlations, and visualizations (box, histograms, density, scatter, maps) to prepare for predicting sale price with random forest.
Explore the Ames data through a comprehensive eda walkthrough that identifies variable types, missingness, distributions, and correlations, and prepares features for a random forest model to predict sale price.
Review how ensemble learning builds a random forest from multiple trees. Explain bagging, feature sampling, tidymodels parallelism, the Ames housing data, and the upcoming EDA and assignment.
Build and compare random forest models to predict Ames housing sale prices using train-test splits, tenfold cross-validation, and hyperparameter tuning, with optional advanced feature filtering and model comparisons.
Walk through the assignment workflow for Tidymodels part 2, building three random forest models with feature selection, cross-validation, and hyperparameter tuning, then compare RMSE.
Explore gradient boosting and the boosting paradigm, learn how XGBoost and Lightgbm fit into tidymodels, and apply them to predicting Ames house prices and regression tasks.
Boosting builds sequential trees that correct prior errors to improve predictions, and includes gradient boosting and AdaBoost, with XGBoost, LightGBM, and CatBoost for scalable tabular data in regression and classification.
Explore XGBoost, the extreme gradient boosting algorithm with regularization, missing-value handling, and fast parallel training for high-accuracy tabular data models.
Explore building xgboost models in tidymodels for regression, focusing on engine setup, hyperparameter tuning with parsnip and dials, and one-hot encoding with step dummy.
Build an XGBoost model in tidymodels for the diamonds dataset, predicting log price and applying one-hot encoding for categorical features. Tune trees, depth, and learning rate via 10-fold cross-validation.
Explore XGBoost in tidymodels part 3, tune hyperparameters such as trees, depth, and learning rate, evaluate with RMSE, visualize results, and finalize the best model for test evaluation.
Use XGBoost with tidymodels to predict Ames house prices, performing data loading, a train-test split, and 10-fold cross-validation with one-hot encoding and RMSE evaluation.
Apply XGBoost to ames house price prediction by performing train/test split, ten-fold cross-validation, and hyperparameter tuning with a recipe that includes one-hot encoding, workflow creation, and model evaluation.
Explore Lightgbm, a fast gradient boosting framework designed for big data, featuring leaf-wise tree growth, native categorical handling, and CPU-based training within tidymodels via bonsai.
Recap the boosting paradigm as a counter to bagging in random forests. Demonstrate building an XGBoost model with tidymodels to predict diamond prices and compare with Lightgbm insights.
Conduct a multi-model assignment on Ames housing data to predict sale price using kNN, trees, random forest, XGBoost, LightGBM, and linear models such as elasticnet, ridge, and lasso.
Explore an assignment walkthrough that tunes multiple tidymodels algorithms (knn, tree, rf, xgboost, lightgbm) with train/validate/test splits, recipes, and workflow sets, selecting the best model via rmse.
Predict Japan real estate transaction prices using a time-aware quarterly dataset (2005–2019) with tidymodels, implementing time-based train/validation/test splits and hyperparameter tuning for XGBoost and LightGBM.
Load and pre-process the train data with tidyverse and janitor, sample 10% for faster exploration, create a date column from year and quarter, and set trade price as the target.
Perform the initial eda with eda functions in tidyverse, inspecting column types, missing data, and per-quarter data points, and filter categorical levels with thresholds to guide preprocessing and imputation decisions.
In data preprocess, the lecturer drops irrelevant columns and columns with too many missing rows to prevent leakage, then parse date and convert categorical columns to factors.
Explore the log-transformed trade price with density and box plots, examine time trends with scatter plots, inspect numeric and categorical distributions, and review correlations for feature engineering.
Master on-the-fly feature engineering for a tidymodels workflow: extract year and quarter, compute years since 2005 and squareness, lump high-cardinality categories, and plan recipes with missing value handling.
Dive into model training with tidymodels, building XGBoost and LightGBM pipelines, using time-based resampling, comprehensive recipes, and hyperparameter tuning for robust predictive models.
Predict real estate prices on the validation set with the final trained model, compare RMSE between Lightgbm and extreme gradient boosting, and highlight Lightgbm as the best.
Test a real estate prediction model using Lightgbm with half a million test records, applying preprocessing and year/quarter features, evaluating with RMSE and plotting predicted versus actual prices.
Explore the Tidymodels part 2 outro as the instructor recaps regression and machine learning context, previews future classification content in part 3 and 4, and outlines release plans.
You've built your first predictive models. You understand linear regression and regularization. Now it’s time to level up.
This course is designed for learners who want to go beyond simple models and tackle non-linear relationships, ensemble algorithms, and real-world modeling challenges with confidence.
What You'll Learn?
In this course, we remain in the regression domain but expand your modeling toolbox with powerful new algorithms and modeling strategies:
Use k-nearest neighbors (KNN) for flexible, non-parametric regression
Build decision trees for interpretable, rule-based models
Apply random forests for robust ensemble modeling
Harness the power of XGBoost and LightGBM, two of the fastest and most powerful tree-based learners
Understand the principles behind bagging and boosting
Learn how parallel processing speeds up model tuning and resampling
Tune hyperparameters efficiently with grids and a Bayesian iterative search approach
Compare models using consistent metrics across algorithms
Structure your modeling workflow for scalability, readability, and reproducibility
And to wrap it all up, you’ll complete a final modeling project, where you build a predictive model on new data, applying everything you’ve learned.
Why Take This Course?
Modern data science requires more than just one-size-fits-all models.
With real-world data, relationships are rarely linear. This course teaches you how to adapt, choose the right model, and justify your choices.
More than just syntax, this course helps you think like a machine learning practitioner while staying fully within the elegant, consistent, and tidy philosophy of tidymodels.
What You’ll Get?
Clear explanations of advanced modeling concepts
Intuitive explanations of ensembles, bagging, and boosting
Step-by-step implementations of each algorithm
Practical coding examples with real data
Exercises and assignments to reinforce learning
Solutions for all exercises and assignments
A final capstone modeling project
All code, datasets, and solutions provided
Lifetime access
Who Is This Course For?
Students who have a basic understanding of tidyverse and tidymodels already (it is strongly recommended to first complete course part 1)
Data analysts and scientists who want to master regression with modern algorithms
R users, who are ready to move beyond linear models into flexible, high-performing learners
Anyone curious about ensemble methods, model tuning, and boosting strategies in R
If you're ready to boost your R modeling skills and learn the most powerful regression tools available today, all inside the tidymodels framework, then this course is for you.
Enroll today and take your predictive modeling to the next level!!!