
Bharani Kumar de Puru introduces himself as an industrial revolution 4.0 implementer and data scientist with 16 years across HSBC, ITC Infotech, Infosys, and Deloitte, plus interlocking director role.
Explore the agenda and stages of analytics and learn how project management methodology provides a 30,000-foot view of real-world analytics projects and module concepts.
Explore diagnostic analytics by asking why events occur, such as spikes and drops in Covid 19 cases, and attributing causes like lockdowns and vaccination to explain trends.
Prescriptive analytics uses what-if analysis and predictions to guide actionable decisions, such as vaccines, awareness campaigns, or automated shutdowns, across descriptive, diagnostic, predictive, and prescriptive stages.
Learn how predictive analytics uses current data to forecast future outcomes, from covid cases to vaccination rates, and how the chosen time horizon affects prediction validity amid changing conditions.
Explore CRISP-ML(Q) framework for data projects by outlining its six phases: business and data understanding, data preparation, model building and tuning, evaluation, deployment, and monitoring and maintenance.
Define the scope of application by articulating the business objective to minimize loan defaulters under constraints, using data and survival analytics to optimize risk and profits.
Define business success criteria by tying KPIs, like loan default under 5%, to the problem; ensure accuracy above 85%, results within one second, and a clear return on investment.
Understand business use cases that set objectives, like minimizing fraud while preserving customer convenience in credit card transactions. Explore drone-driven precision farming balances yield with cost.
Explore agenda data understanding by identifying data types, scales of measurement, key terms, and primary and secondary data collection techniques.
Explore data understanding by defining data as measurable, analyzing it to build models for predictions and optimization, and applying what-if analysis to support management decisions under constraints.
Identify continuous versus discrete data using decimal representations, with examples like time, money, height, weight, and discrete examples such as laptops or cars; touch on count data and categorical data.
Explore categorical data versus count data and their subtypes, including binary, multiple categorical, and count data, with real-world examples like churn, defaults, and claims, alongside continuous data concepts.
Examine real-world data understanding by distinguishing nominal, ordinal, interval, and ratio data with travel examples, including flight numbers, gate order, temperatures, and money with absolute zero.
Explain the scale of measurement from nominal to ratio data, and show what counts, proportions, and percentages you can compute, highlighting ratio data as enabling the most comprehensive analysis.
Compare quantitative and qualitative data, including structured versus unstructured data, defining numerical, continuous, and count data as quantitative, and categorical data as qualitative, to illuminate which data guides decision making.
Explain how structured data fits in tabular formats, while unstructured data such as videos, images, audio, and textual data requires transformation to structured formats.
Define data collection and distinguish primary from secondary data sources; identify outputs as response or dependent variables and inputs as explanatory or independent variables; convert to structured format for analysis.
Understand primary data sources through real-world examples, including using social media data and sentiment analysis to augment bank data, and distinguishing primary data from secondary data captured by IoT sensors.
Understand secondary versus primary data sources, including internal, external, open source, and syndicated data, and learn to combine drone analytics and map data to inform hypothesis testing in market planning.
Explore how to design surveys for data collection by translating business reality into root-cause analysis, decision problems, and clear research objectives, then decompose into one-dimensional constructs and craft targeted questions.
Learn how design of experiments guides data collection and marketing decisions by testing coupon discounts, expiry dates, and customer radius to reveal patterns in redemption.
Identify random and systematic data collection errors, apply gauge R&R and attribute agreement analysis, enforce standard operating procedures, and recognize biases to improve data quality and generalizability.
Address bias and fairness by using a diverse dataset and avoiding race or gender-based variables, while recognizing data understanding and preparation are essential.
Explore the crisp-ml(q) data preparation framework, with six phases from business and data understanding to data collection, recording objectives, constraints, and consider secondary data sources or primary data collection.
Master the probability formula, the ratio of interested events to the total number of events. Apply it to dice, with examples like greater than three and smaller than four.
Explore how a random variable blends a variable output with associated probabilities, illustrated by coin flips and dice, and how probability distributions sum to one.
Explore probability theory and its application, including probability distributions, random variables, and the discrete versus continuous distinction, illustrated with iPad sales data.
Explore the box plot and the differences between percentile, quantile, and quartile, including Q1, Q2, Q3, and the minimum value and maximum value, and their relationships.
Explore the hierarchical clustering process using a five-point toy dataset, computing Euclidean distances, applying single linkage, and visualizing merges in a dendrogram.
Explore hypothesis testing as inferential statistics, connecting population parameters to sample conclusions; compare parametric and nonparametric tests, and learn about null and alternate hypotheses, p-values.
Explore business use cases of hypothesis testing, define null and alternate hypotheses, and decide actions like digital marketing or hiring using a normality to nonparametric testing flow.
Apply a two-sample t test to compare two promotions on purchases, with purchases as output and promotions as discrete inputs. Conduct normality testing and specify null and alternative hypotheses.
Apply one-way anova to compare transaction times across three suppliers, verify normality and equal variances, and interpret f-statistics and p-values to decide if all means are equal.
Explore confidence intervals for normally distributed data using z transformations and the standard normal distribution. Learn to compute probabilities and p values with z tables and SciPy.
This course provides a comprehensive introduction to the fundamental concept of hypothesis testing in statistics. Hypothesis testing is a critical tool for making informed decisions and drawing meaningful conclusions from data. Through a combination of theoretical concepts and practical applications, students will learn how to formulate hypotheses, perform hypothesis tests, interpret results, and make valid inferences about populations based on sample data.
Course Objectives:
By the end of this course, students should be able to:
Understand the purpose and importance of hypothesis testing in various fields.
Differentiate between null and alternative hypotheses and select appropriate test criteria.
Apply various hypothesis testing methods for means, proportions, and variances.
Interpret p-values, confidence intervals, and effect sizes to make informed conclusions.
Determine sample sizes for hypothesis tests and assess the power of tests.
Identify and mitigate common errors and misconceptions in hypothesis testing.
Course Outline:
Introduction to Hypothesis Testing
Role of hypothesis testing in data analysis
Formulating null and alternative hypotheses
Significance level and p-values
Probability and Distributions Review
Probability distributions and their properties
Sampling distributions and central limit theorem
Hypothesis Testing Process
Steps in hypothesis testing
One-tailed vs. two-tailed tests
One-Sample Hypothesis Tests
Z-tests and t-tests for means and proportions
Interpreting results and drawing conclusions
Two-Sample Hypothesis Tests
Independent sample tests and paired sample tests
Comparing means and proportions
Analysis of Variance (ANOVA)
One-way ANOVA for multiple group comparisons
Post hoc tests and multiple comparisons
Non-Parametric Tests
Introduction to non-parametric tests
When to use non-parametric methods.
Note to Students:
This course is designed to provide you with essential tools for drawing meaningful conclusions from data. Engaging in class activities, seeking help when needed, and actively participating will contribute to a successful learning experience.
Course Duration and Format:
This is a [semester/quarter]-long course, consisting of [number] of weekly sessions. Each session will typically last [duration] and will involve a mix of lectures, discussions, and practical exercises. Additionally, there may be optional review sessions or office hours to provide extra support for students.
Course Learning Outcomes:
By the end of this course, students will be able to:
Formulate clear null and alternative hypotheses for different research questions.
Choose appropriate hypothesis testing methods based on data type and study design.
Perform hypothesis tests using statistical software and interpret the results.
Evaluate the significance of p-values and make informed decisions based on them.
Calculate and interpret confidence intervals to estimate population parameters.
Additional Resources:
In addition to the core course materials, students will have access to supplementary resources, including:
Recommended readings and articles for deeper understanding.
Online tutorials and video demonstrations of hypothesis testing procedures.
Sample datasets for practice and exploration outside of class.
Reference guides on statistical software usage.
This course provides a comprehensive exploration of hypothesis testing, empowering students with the skills to analyze data, draw meaningful conclusions, and contribute to evidence-based decision-making across various fields. Through a combination of theoretical knowledge, practical exercises, and real-world applications, students will develop a solid foundation in statistical inference, setting them on a path to becoming proficient data analysts and informed researchers.