
Learn to build data science models from data and algorithms using Python, with real-world examples like Netflix recommendations and fraud detection, covering data analysis, preprocessing, exploratory data analysis, and modeling.
Explore how data science teams organize around business problems, with data scientists collaborating with business analysts, data engineers, architects, and production engineers to deploy and monitor models across the lifecycle.
Explore data management with numpy and pandas, visualize with matplotlib and seaborn, and apply analytics and machine learning in Python, distinguishing data science from analytics.
Explore the data science and machine learning relationship, including data collection, cleaning, analysis, visualization, and predictive modeling, and learn the CRISP-DM lifecycle from business understanding to deployment.
Explore descriptive and inferential statistics, probability theories, and hypothesis testing for data analysis. Learn about randomized controlled experiments, double-blind designs, confounding factors, and the distinction between association and causation.
Learn how histograms reveal the normal distribution across income and spending data, use class intervals and frequencies to build a standard unit curve, and perform normal approximation with 68–95–99.7 guidance.
Explore the central limit theorem by drawing samples from any population, showing the distribution of sample means tends toward normal, with implications for sampling, standard deviation, and hypothesis testing.
Master probability theory through practical chance models, from dice outcomes to ball drawing scenarios, and apply multiplication and addition rules to compute event likelihoods.
Explore probability with the addition rule for mutually exclusive events, and apply the binomial formula to compute exact outcomes. Learn about expected value and standard error across practical examples.
Learn hypothesis testing, including framing null and alternative hypotheses, calculating z-test statistics, standard error, and p-values to assess if observed differences reflect real population effects.
This attachment contains the code and data for the entire course.
Launch Python via Anaconda and work in a Jupyter notebook, learning modular code cells and pandas data frames for data management, while comparing Jupyter with Spider and using OS commands.
Learn definite and indefinite iterations in python through while and for loops, processing lists, tuples, sets, dictionaries, and strings, with increment and decrement operators and key-based value access.
Learn how Python iterations work with ranges and for loops, including two- and three-parameter ranges, custom iterators with yield, and controlling loops with break and continue.
Explore Python lists and other collection types by learning creation, indexing, slicing, mutability, and common operations like append, extend, range operations, popping, and removing elements.
Learn how Python tuples provide immutable sequences, create nested structures from lists and dictionaries, and perform indexing, slicing, and common operations like enumerate, zip, any, all, max, and sum.
Explore Python dictionaries, their curly braces syntax, key–value pairs, and nested structures; learn creation, access, insertion and update, deletion, and safe retrieval with get and try/except.
Explore Python dictionaries by using items, keys, and values to access data; learn pop item and setdefault behaviors, and iterate with for loops to print keys and values.
Explore Python sets: unordered, unique elements. Learn unions, intersections, differences, and symmetric differences, plus membership tests, add, update, remove, copy, and subset checks, with iteration.
Learn that assignment creates a shared reference to a set, while using the copy function yields a physically separate copy you can verify by updating one set.
Explore how numpy handles numeric data as multi-dimensional arrays, create arrays with zeros, ones, full, identity, random values, arange sequences, and reshape, transpose, and dot products.
Explore numpy arrays operations including dot product, transpose, flatten, and copy, and learn to round, convert types, and access or slice array sections to extract cross-sections.
Explore numpy arrays and their 2d iteration by rows and columns, practice reshaping with range, and master stacking, splitting, and copying arrays using vstack, hstack, vsplit, split, and copy.
Explore pandas series and data frames by creating one-dimensional series from lists or dictionaries, manage indexes, handle missing data, and enable interoperation with numpy and common file formats.
Master pandas series indexing with loc and iloc to slice, pull single elements, and inspect index, values, and data types. Create date_range indices and retrieve time-based values with at_time.
Explore Pandas series operations in practical data science using Python, including between for range tests, shallow and deep copies, descriptive statistics with describe, and drop, fill, replace.
Explore how the describe function computes statistics for numeric columns in a data frame, including count, mean, standard deviation, min, percentiles, and max, plus sorting, transposing, and conditional filtering.
Explore practical data frame manipulation in pandas, including adding records with series, updating cells, creating derived columns, resetting indices, and concatenating data frames in horizontal and vertical axes.
Learn pandas dataframe merging with keys, performing inner, left, right, and outer joins to apply country codes and align Gap Minder records across years.
Apply smart missing value replacements by column mean or fixed values, then use groupby and pivot tables to compute means, counts, and other aggregations by continent and country.
Explore how to define and use Python user-defined functions, pass multiple input types, return multiple values, and manage global and local scope.
Learn python lambda functions—anonymous, single-expression tools that can be faster than regular functions. Apply them with map to dataframe life expectancy fields and with reduce to fold values.
Explore practical Python string operations, including creating strings with quotes, indexing and slicing, concatenation, iteration, membership checks, and formatting, joining, splitting, replacing, and case conversion.
Entire code and data are attached as resource with the first lecture of Python section.
Use exploratory data analysis to uncover insights and prepare data for machine learning, applying univariate and multivariate methods, charts, and feature engineering on Lending Club data to identify default drivers.
Perform exploratory data analysis on the Lending Club dataset to identify features linked to defaults, and engineer a defaulted indicator from loan status, for clearer insights.
Explore loan amount distributions using box plots and quantiles, compare defaulted and non-defaulted groups, and uncover univariate insights on annual income, loan, funded amounts, and interest rates.
Engage in an eda project on loan status, using a bar plot for charged off, fully paid, and current loans, with a heatmap and null-value checks via info and isna.
Identify fields with high missing value percentages using a dataframe function, assess categorical feature distributions, and build income bins to explore default ratios with bar plots.
Bin funded amount into categorical ranges, combine with annual income categories, and visualize default ratios with seaborn to extract insights on loan performance and data relationships.
Analyzes how debt-to-income ratio, employment tenure, loan purpose, and interest rates influence loan defaults using box plots, bins, and cross-tab analyses to guide LendingClub decisions.
Explore how machine learning uses data to train algorithms that build models for predictions and decisions, including voice assistants, stock price forecasts, email filtering, recommendations, predictive maintenance, and reinforcement learning.
Trace the history of machine learning from early neural concepts and the Turing test to the GPU-driven era of deep learning, CNNs, AlexNet, and cloud GPUs.
Explore how machine learning solves problems traditional programming cannot, with spam filtering, image and voice processing, fraud detection, and core types: supervised, unsupervised, semi supervised, and reinforcement learning.
Learn how data quality and careful data pre-processing shape machine learning models, with emphasis on selecting and transforming numerical, categorical, time series, textual, and image data for accurate predictions.
Explore linear regression and multivariate regression for predicting continuous outputs like house price, using gradient descent to minimize error and compare models with MSE, RMSE, MAE, RSS, TSS, and R-squared.
Explore classification models that map inputs to discrete outputs and use the confusion matrix to evaluate accuracy, recall, specificity, precision, PPV, and F1 score.
Compare accuracy with precision, recall, and specificity using a heart condition example to show when accuracy can mislead. Learn how F1 score guides balancing false negatives and false positives.
Master linear regression to model the relationship between predictor variables and a target, using simple and multiple regression, cost functions, and data preprocessing on price prediction with Carsales data.
Learn how linear regression models relationships between predictor variables and a target variable, train model parameters, and minimize cost function via mean squared error to obtain a best fit line.
Build a practical linear regression model to predict car resale prices from vehicle attributes, using historical data, preprocessing, feature selection, and evaluation on unseen data.
Apply feature scaling via normalization and standardization to numeric features before linear regression to mitigate gradient descent and distance-based algorithm issues; split data for training and testing to validate generalization.
Explore linear regression with OLS assumptions, test residual normality and homoscedasticity, assess multicollinearity via VIF, and use adjusted R-squared, p-values, and feature elimination in a Python car price prediction project.
Apply linear regression to predict car prices from historical attributes, after exploring data and preprocessing steps that convert categorical features to numeric and clean anomalies.
Identify linear relationships and multicollinearity in the car dataset using pair plots and correlation analysis, drop highly correlated features, and convert ordinal categoricals to dummy variables for stable linear regression.
Split the data into training and test sets, scale features, and build a linear regression model using ordinary least squares to predict car prices, then evaluate with R-squared and p-values.
Evaluate linear regression by plotting vs predicted values, assess residual normality with a probability plot and residuals vs fitted values, achieving 92% training and 86.5% test adjusted R-squared on Carsales.
Explore logistic regression and the logit model, linking predictors to odds via a log-linear equation and a sigmoid probability output, with beta coefficients and gradient descent-based maximum likelihood optimization.
Use logistic regression to predict telecom churn by merging churn, customer, and internet data on customer ID, examining monthly charges, tenure, and contract types.
explore logistic regression for churn prediction through feature engineering, binary to numeric conversion, one-hot encoding of nominal categories, train-test split, and standard scaling.
Guide building a logistic model by cleaning data, checking correlations with a correlation matrix, adding a constant for intercept, and selecting features via p-values below 0.05.
Explore building a logistic regression model with scikit-learn, noting L2 regularization and default intercept. Evaluate performance using a confusion matrix, accuracy, precision, recall, specificity, and roc auc.
Explore logistic regression model optimization by examining how varying probability thresholds affect accuracy, sensitivity, and specificity, using confusion matrices and threshold plots to identify an optimal cutoff.
Explore logistic regression model optimization for churn prediction, balancing accuracy, sensitivity, specificity, and precision-recall tradeoffs to determine optimum probability thresholds and validate on scaled test data.
Explore a worked example of k-means clustering on customer data, with gender, age, income, and spending score, using inertia and silhouette score to determine six clusters.
Explore optimizing k-means clustering by assigning labels, inspecting cluster centers with pivot tables, and visualizing clusters for customer segmentation and promotions, using PCA for dimensionality reduction when needed.
Learn the naive Bayes classifier, a fast probabilistic method based on Bayes theorem and independence assumptions, with Gaussian and other variants, applied to spam filtering, sentiment analysis, and disease detection.
Build and evaluate a gaussian naive bayes classifier using train-test split, train on features X and target y, and assess performance with a confusion matrix and accuracy.
Explore how a decision tree turns data into a rule-based, tree-structured model by selecting the best feature and threshold via Gini impurity or entropy, from root to leaf with pruning.
Tune decision trees with hyperparameters such as min samples split, min samples leaf, max depth, and max features. Understand impurity measures like Gini index and entropy, with iris dataset example.
Tune decision tree hyperparameters with grid search cross-validation, optimizing max depth, min samples split/leaf, max features, and criterion (entropy or gini) on the iris dataset for high accuracy.
Are you aspiring to become a Data Scientist or Machine Learning Engineer? if yes, then this course is for you.
In this course, you will learn about core concepts of Data Science, Exploratory Data Analysis, Statistical Methods, role of Data, Python Language, challenges of Bias, Variance and Overfitting, choosing the right Performance Metrics, Model Evaluation Techniques, Model Optmization using Hyperparameter Tuning and Grid Search Cross Validation techniques, etc.
You will learn how to perform detailed Data Analysis using Pythin, Statistical Techniques, Exploratory Data Analysis, using various Predictive Modelling Techniques such as a range of Classification Algorithms, Regression Models and Clustering Models. You will learn the scenarios and use cases of deploying Predictive models.
This course covers Python for Data Science and Machine Learning in great detail and is absolutely essential for the beginner in Python.
Most of this course is hands-on, through completely worked out projects and examples taking you through the Exploratory Data Analysis, Model development, Model Optimization and Model Evaluation techniques.
This course covers the use of Numpy and Pandas Libraries extensively for teaching Exploratory Data Analysis. In addition, it also covers Marplotlib and Seaborn Libraries for creating Visualizations.
There is also an introductory lesson included on Deep Neural Networks with a worked-out example on Image Classification using TensorFlow and Keras.
And in the last section, you will learn how to create a FAST API using your ML Model just as you need to deploy your Model in production, and invoke the FAST API from a Streamlit UI.
Course Sections:
Introduction to Data Science
Use Cases and Methodologies
Role of Data in Data Science
Statistical Methods
Exploratory Data Analysis (EDA)
Understanding the process of Training or Learning
Understanding Validation and Testing
Python Language in Detail
Setting up your DS/ML Development Environment
Python internal Data Structures
Python Language Elements
Pandas Data Structure – Series and DataFrames
Exploratory Data Analysis (EDA)
Learning Linear Regression Model using the House Price Prediction case study
Learning Logistic Model using the Credit Card Fraud Detection case study
Evaluating your model performance
Fine Tuning your model
Hyperparameter Tuning for Optimising our Models
Cross-Validation Technique
Learning SVM through an Image Classification project
Understanding Decision Trees
Understanding Ensemble Techniques using Random Forest
Dimensionality Reduction using PCA
K-Means Clustering with Customer Segmentation
Introduction to Deep Learning
Bonus Module: Time Series Prediction using ARIMA
Building a FAST API to deploy your ML Model