
Explore linear and logistic regression with R Studio, learn their math concepts and implement models in R, including data preprocessing, fitting, and evaluating results for business insights.
Identify the two main data types: qualitative and quantitative, and distinguish their subtypes, such as nominal, ordinal, discrete, and continuous, to guide appropriate analyses.
Celebrate reaching this milestone and stay motivated to complete the course, as you join the top 50% of learners, rate the course, and access Q&A, AI assistant, and a certificate.
Explore descriptive and inferential statistics, including measures of center and dispersion, frequency distributions, bar charts, and histograms, with neural networks covered as a key deep learning technique.
Explore frequency distributions, bar charts, histograms, and relative frequency; compare qualitative and quantitative data and grasp normal distribution, skewness, and uniformity.
Explore descriptive measures of center in data, including mean, median, mode, and mid-range, with examples and distinctions between population mean and sample mean.
Explore measures of dispersion, including range, standard deviation, and variance, and how variance equals the square of the standard deviation. Note that the range is influenced by outliers.
Install R first, then R Studio on Windows, then complete a quick crash course in statistics using the R Studio interface and scripts.
Learn basic R and R studio tasks: write and run commands, comment with #, assign with <-, create vectors with c() or 1:10, manage workspace with ls and rm.
Learn to install, load, and manage packages in R, using lib linear for linear regression and ggplot for plotting. Explore built-in packages and repository installations for reproducible code.
Identify and load built-in data sets in R using the data sets package, inspect iris with str and help, and bring the data into the workspace for analysis.
Explore manual data entry in R: assign values to variables using direct, concatenation, and sequence methods, and input values with the scan function.
Import data from tab-delimited text and comma-delimited csv into R Studio using read.table and read.csv, with headers, revealing product (1862 obs, 4 vars) and customer (793 obs, 9 vars) datasets.
Create bar plots in R to visualize region-based frequency distributions, order and orient bars, customize color and borders, add titles and axis labels, and export images for presentations.
Create a histogram in R Studio using hist to plot age distributions, customize breaks and categories, set frequency, color, labels, and export the chart.
Identify the business context and the factors that impact the variables of interest. Gather primary and secondary research to define variables like cart abandonment and choose a model with data.
Explore data by identifying internal and external data needs, requesting data from stakeholders, and performing quality checks. Use cart abandonment and channel insights to define variables and tidy the data.
Examine a real estate pricing dataset: 506 observations across 19 variables, with a data dictionary and primary key, and learn how to join sources and define variables for regression analysis.
Import the house pricing dataset into R Studio with df <- read.csv and header = true, then view df and str(df) to see 506 observations across 19 variables.
Explore univariate analysis using descriptive statistics to summarize single variables, including mean, median, mode, dispersion and quartiles, and learn how the Extended Data Dictionary reveals outliers and missing values.
Perform univariate analysis in R using the extended data dictionary, plot histograms and bar charts, and identify outliers, skewness, missing values, and useless variables for modeling.
Identify and treat outliers in data using box plots, scatter plots, and histograms. Apply capping and floating with percentile-based or sigma methods to improve prediction accuracy in regression models.
Cap outliers in n_hot_rooms at three times the 99th percentile and floor rainfall at 0.3 times the first quantile using quantile in R, improving mean and median alignment.
Learn how to handle missing values in data for regression by deciding to delete observations or impute with values like mean, median, mode, or zero when sensible, plus segment means.
Impute missing values in R by replacing NA with the mean, using na.rm=TRUE, identify NA positions with is.na, apply the substitution, and verify no NA remains.
Explore seasonality in time-based data and learn how to remove it by computing month factors from yearly and monthly means, then normalize sales for better model fit.
Examine bivariate analysis with scatterplots and correlation matrices to assess two-variable relationships and decide on linearity or transformations, and apply variable transformations to improve model fit and handle multicollinearity.
Transform the crime rate with a log (add one) to achieve a linear relationship with price, and create average distance from dist1–dist4, removing bus terminal after EDD.
Apply univariate and bi variate analysis to keep variables, removing single-value ones like bus terminal. Consider regulatory constraints; use business knowledge and scatterplots to confirm relationships and iterate post-regression.
Create dummy variables to convert non-numeric categorical data, such as airport yes/no and favorite subjects, into 0/1 indicators for regression, using n minus one variables to avoid implying order.
Create dummy variables in R using the dummies package to convert categorical fields like airport and water body into numerical indicators for regression, keeping one fewer dummy than categories.
Explore correlation matrices and coefficients to identify positive, negative, and near-zero relationships, distinguish correlation from causation, and address multicollinearity by selecting the most informative variables.
Use the correlation matrix in R to identify price-related variables, round decimals for readability, and remove parks due to high correlation with air quality to avoid multi collinearity.
Explore how linear regression serves as a simple, foundational tool for supervised learning. Learn the least squares approach and apply it to house price prediction to estimate variable effects.
Apply simple linear regression with one predictor to model y as beta0 plus beta1 x, estimate beta0 cap and beta1 cap via least squares, and interpret residuals and RSS.
Assess the accuracy of sample beta0 and beta1 using residual standard error and confidence intervals to test for a relationship between house price and number of rooms.
Assess model accuracy by comparing predicted and actual values with residual standard error and R squared, examine TSS and RSS, and consider adjusted R squared for model complexity.
Learn to run a simple linear regression in R using lm, with price as dependent variable and room_num as independent variable from df, and interpret beta, p-value, and R-squared.
Extend simple regression to multiple linear regression with 16 predictors, interpreting each beta as the effect of a unit change while others fixed, covering RSS, R-squared, and p-values.
Use the f statistic in multiple regression to test whether the model’s predictors jointly relate to the response, with a 5% p-value threshold to guard against false significance.
Transform categorical variables into dummy variables in linear regression, and interpret coefficients and p-values to assess impact on house price; airport and water body illustrate baseline contrasts.
Learn how to run a multiple linear regression in R using lm with multiple predictors, and interpret coefficients, p-values, r-squared and adjusted r-squared for house price insights.
Split data into training and test sets to evaluate training error and test mean squared error, minimize test error, and guide model selection with validation set, leave-one-out, and k-fold cross-validation.
Explain the bias-variance trade-off in regression, showing how increasing model flexibility lowers bias but raises variance, and how the total test error is minimized at an optimal balance.
Split data into train and test sets using ca tools in R, train a linear model on the training data, and compare mean squared error on train and test data.
Learn three lightweight classification models—logistic regression, k-nearest neighbours, and linear discriminant analysis (LDA)—for predicting binary outcomes, using a preprocessed dataset of 506 property transactions where sold equals 0 or 1.
Import data into a variable named df with read.csv and header = TRUE, adjust backslashes to forward slashes, then view the structure with str(df) to see variables and data types.
Explore two business questions in regression: prediction and inferential, using house data to predict sale within three months and to estimate each variable's impact on the response.
Explore why linear regression cannot be used for classification—issues with multi-level responses, dummy variables, probability interpretation, and sensitivity to outliers—then see how logistic regression overcomes them.
Explore logistic regression for predicting default probability using a sigmoid function, with balance and income as predictors, and learn maximum likelihood estimation to fit the model on credit default data.
Train a simple logistic model in R with one predictor, interpret beta0 and beta1, and assess significance via standard error, z value, and p-value against a threshold.
Using a one-predictor logistic model, we estimate beta0 = 6.1 and beta1 = -0.03, and a p-value of 0.0006 shows price impacts the response.
Extend logistic regression to multiple predictors, estimate all betas by maximum likelihood, and compute probabilities to classify outcomes using a 0.5 boundary.
Build a logistic regression with multiple predictors in Python using sklearn and statsmodels, using all features except the sold variable as X and sold as y; fit and inspect coefficients.
Learn to use a confusion matrix to compare true versus predicted classes and identify type one and type two errors—false positives and false negatives in R Studio.
Analyze classifier performance using confusion matrices, exploring true positives, false positives, true negatives, and false negatives, precision, specificity, and sensitivity. Learn to interpret ROC curves and AUC to compare models.
Compute predicted probabilities from the glm model with predict(type='response'), assign yes/no classes using a 0.5 threshold, and evaluate the results with a confusion matrix.
Split data into training and test sets to evaluate model accuracy on unseen data using a confusion matrix and validation set, leave-one-out, and k-fold cross-validation.
Split data into 80% training and 20% test with a seed, train logistic regression, and evaluate with a confusion matrix on the test set.
Celebrate reaching the final milestone in the linear regression and logistic regression course using R Studio and receive a certificate of completion by email.
You're looking for a complete Linear Regression and Logistic Regression course that teaches you everything you need to create a Linear or Logistic Regression model in R Studio, right?
You've found the right Linear Regression course!
After completing this course you will be able to:
Identify the business problem which can be solved using linear and logistic regression technique of Machine Learning.
Create a linear regression and logistic regression model in R Studio and analyze its result.
Confidently practice, discuss and understand Machine Learning concepts
A Verifiable Certificate of Completion is presented to all students who undertake this Machine learning basics course.
How this course will help you?
If you are a business manager or an executive, or a student who wants to learn and apply machine learning in Real world problems of business, this course will give you a solid base for that by teaching you the most popular technique of machine learning, which is Linear Regression
Why should you choose this course?
This course covers all the steps that one should take while solving a business problem through linear regression.
Most courses only focus on teaching how to run the analysis but we believe that what happens before and after running analysis is even more important i.e. before running analysis it is very important that you have the right data and do some pre-processing on it. And after running analysis, you should be able to judge how good your model is and interpret the results to actually be able to help your business.
What makes us qualified to teach you?
The course is taught by Abhishek and Pukhraj. As managers in Global Analytics Consulting firm, we have helped businesses solve their business problem using machine learning techniques and we have used our experience to include the practical aspects of data analysis in this course
We are also the creators of some of the most popular online courses - with over 150,000 enrollments and thousands of 5-star reviews like these ones:
This is very good, i love the fact the all explanation given can be understood by a layman - Joshua
Thank you Author for this wonderful course. You are the best and this course is worth any price. - Daisy
Our Promise
Teaching our students is our job and we are committed to it. If you have any questions about the course content, practice sheet or anything related to any topic, you can always post a question in the course or send us a direct message.
Download Practice files, take Quizzes, and complete Assignments
With each lecture, there are class notes attached for you to follow along. You can also take quizzes to check your understanding of concepts. Each section contains a practice assignment for you to practically implement your learning.
What is covered in this course?
This course teaches you all the steps of creating a Linear Regression model, which is the most popular Machine Learning model, to solve business problems.
Below are the course contents of this course on Linear Regression:
Section 1 - Basics of Statistics
This section is divided into five different lectures starting from types of data then types of statistics
then graphical representations to describe the data and then a lecture on measures of center like mean
median and mode and lastly measures of dispersion like range and standard deviation
Section 2 - Python basic
This section gets you started with Python.
This section will help you set up the python and Jupyter environment on your system and it'll teach
you how to perform some basic operations in Python. We will understand the importance of different libraries such as Numpy, Pandas & Seaborn.
Section 3 - Introduction to Machine Learning
In this section we will learn - What does Machine Learning mean. What are the meanings or different terms associated with machine learning? You will see some examples so that you understand what machine learning actually is. It also contains steps involved in building a machine learning model, not just linear models, any machine learning model.
Section 4 - Data Preprocessing
In this section you will learn what actions you need to take a step by step to get the data and then
prepare it for the analysis these steps are very important.
We start with understanding the importance of business knowledge then we will see how to do data exploration. We learn how to do uni-variate analysis and bi-variate analysis then we cover topics like outlier treatment, missing value imputation, variable transformation and correlation.
Section 5 - Regression Model
This section starts with simple linear regression and then covers multiple linear regression.
We have covered the basic theory behind each concept without getting too mathematical about it so that you
understand where the concept is coming from and how it is important. But even if you don't understand
it, it will be okay as long as you learn how to run and interpret the result as taught in the practical lectures.
We also look at how to quantify models accuracy, what is the meaning of F statistic, how categorical variables in the independent variables dataset are interpreted in the results, what are other variations to the ordinary least squared method and how do we finally interpret the result to find out the answer to a business problem.
By the end of this course, your confidence in creating a regression model in Python will soar. You'll have a thorough understanding of how to use regression modelling to create predictive models and solve business problems.
Go ahead and click the enroll button, and I'll see you in lesson 1!
Cheers
Start-Tech Academy
------------
Below is a list of popular FAQs of students who want to start their Machine learning journey-
What is Machine Learning?
Machine Learning is a field of computer science which gives the computer the ability to learn without being explicitly programmed. It is a branch of artificial intelligence based on the idea that systems can learn from data, identify patterns and make decisions with minimal human intervention.
What is the Linear regression technique of Machine learning?
Linear Regression is a simple machine learning model for regression problems, i.e., when the target variable is a real value.
Linear regression is a linear model, e.g. a model that assumes a linear relationship between the input variables (x) and the single output variable (y). More specifically, that y can be calculated from a linear combination of the input variables (x).
When there is a single input variable (x), the method is referred to as simple linear regression.
When there are multiple input variables, the method is known as multiple linear regression.
Why learn Linear regression technique of Machine learning?
There are four reasons to learn Linear regression technique of Machine learning:
1. Linear Regression is the most popular machine learning technique
2. Linear Regression has fairly good prediction accuracy
3. Linear Regression is simple to implement and easy to interpret
4. It gives you a firm base to start learning other advanced techniques of Machine Learning
How much time does it take to learn Linear regression technique of machine learning?
Linear Regression is easy but no one can determine the learning time it takes. It totally depends on you. The method we adopted to help you learn Linear regression starts from the basics and takes you to advanced level within hours. You can follow the same, but remember you can learn nothing without practicing it. Practice is the only way to remember whatever you have learnt. Therefore, we have also provided you with another data set to work on as a separate project of Linear regression.
What are the steps I should follow to be able to build a Machine Learning model?
You can divide your learning process into 4 parts:
Statistics and Probability - Implementing Machine learning techniques require basic knowledge of Statistics and probability concepts. Second section of the course covers this part.
Understanding of Machine learning - Fourth section helps you understand the terms and concepts associated with Machine learning and gives you the steps to be followed to build a machine learning model
Programming Experience - A significant part of machine learning is programming. Python and R clearly stand out to be the leaders in the recent days. Third section will help you set up the Python environment and teach you some basic operations. In later sections there is a video on how to implement each concept taught in theory lecture in Python
Understanding of Linear Regression modelling - Having a good knowledge of Linear Regression gives you a solid understanding of how machine learning works. Even though Linear regression is the simplest technique of Machine learning, it is still the most popular one with fairly good prediction ability. Fifth and sixth section cover Linear regression topic end-to-end and with each theory lecture comes a corresponding practical lecture where we actually run each query with you.
Why use Python for data Machine Learning?
Understanding Python is one of the valuable skills needed for a career in Machine Learning.
Though it hasn’t always been, Python is the programming language of choice for data science. Here’s a brief history:
In 2016, it overtook R on Kaggle, the premier platform for data science competitions.
In 2017, it overtook R on KDNuggets’s annual poll of data scientists’ most used tools.
In 2018, 66% of data scientists reported using Python daily, making it the number one tool for analytics professionals.
Machine Learning experts expect this trend to continue with increasing development in the Python ecosystem. And while your journey to learn Python programming may be just beginning, it’s nice to know that employment opportunities are abundant (and growing) as well.
What is the difference between Data Mining, Machine Learning, and Deep Learning?
Put simply, machine learning and data mining use the same algorithms and techniques as data mining, except the kinds of predictions vary. While data mining discovers previously unknown patterns and knowledge, machine learning reproduces known patterns and knowledge—and further automatically applies that information to data, decision-making, and actions.
Deep learning, on the other hand, uses advanced computing power and special types of neural networks and applies them to large amounts of data to learn, understand, and identify complicated patterns. Automatic language translation and medical diagnoses are examples of deep learning.