
Explore biostatistics with Python to analyze infectious disease data, visualize trends with heatmap and scatter plots, forecast outbreaks, and inform public health policy.
Learn to analyze infectious disease data using Python for biostatistics, covering tools, datasets, time series decomposition, cleaning, exploratory analysis, and forecasting with pandas, NumPy, matplotlib, and scikit-learn in Google Colab.
Identify target audiences—biostatisticians and epidemiologists, public health specialists, and data scientists—and learn Python-based biostatistics for infectious disease analysis, epidemiology modeling, forecasting, and data visualization to inform data-driven decisions.
Explore setting up Python with pandas, matplotlib, numpy, and scikit learn, and using VS Code, Google Colab, or Jupyter to access Kaggle datasets and dataset search for infectious disease analysis.
Explore biostatistics fundamentals and apply time series and epidemiological modeling to infectious disease data, evaluate clinical trial and genomic analyses, build risk models, and address data privacy and quality challenges.
Learn to calculate infectious disease transmissions using a three-equation seir model with a 1000-person population, then compute new infections, recoveries, susceptibles, and derive the basic reproduction number R0.
Analyze six factors that accelerate the spread of infectious disease, including population density, travel and mobility, environment, health care accessibility, herd immunity, and antigenic variation.
Learn to set up Google Colab as a browser-based Python IDE, log in with Gmail, and run code blocks for epidemiological modeling and infectious disease data forecasting.
Explore Kaggle to locate infectious disease datasets, filter by category, inspect a dataset with a 2 MB size, then download and unzip it for analysis.
Upload and read infectious disease datasets in Google Colab using pandas and read_csv, guiding the upload workflow with Google Colab files and correct file naming.
Explore infectious disease datasets by inspecting the first rows, understanding the data structure, shape and data types across ten columns and 141,777 rows with object, integer, and float types.
Learn how to clean infectious disease data by checking and removing missing values and duplicates, then save the cleaned dataset to csv using dropna and drop_duplicate with inplace true.
Learn to detect potential outliers in biostatistics data using z-score thresholds, focusing on count and rate columns, with numpy for calculations.
Welcome to Python for Biostatistics: Analyzing Infectious Diseases Data course. This is a comprehensive project-based course where you will learn step by step on how to perform complex analysis and visualization on infectious diseases datasets. This course is a perfect combination between biostatistics and Python, equipping you with the tools and techniques to tackle real-world challenges in public health. The course will be mainly concentrating on three major aspects, the first one is data analysis where you will explore the infectious diseases data from multiple perspectives, the second one is time series forecasting where you will be guided step by step on how to forecast the spread of infectious diseases using STL model, and the third one is public health policy where you will learn how to make a data driven public health policy based on epidemiological modeling. In the introduction session, you will learn the basic fundamentals of biostatistics, such as getting to know more about challenges that we commonly face when analyzing biostatistics data and statistical models that we will use, for instance STL which stands for seasonal trend decomposition. Then, you will continue by learning how to calculate infectious disease transmission using Kermack-McKendrick equation, this is a very important concept that you need to understand before getting into the coding session. Afterward, you will also learn several factors that can potentially accelerate the spread of infectious diseases, such as population density, healthcare accessibility, and antigenic variation. Once you have learnt all necessary information about biostatistics, we will start the project. Firstly, you will be guided step by step on how to set up Google Colab IDE. Not only that, you will also learn how to find and download infectious diseases dataset from Kaggle. Once, everything is ready, we will enter the main section of the course which is the project section The project will be consisted of three main parts, the first part is to conduct exploratory data analysis, the second part is to build forecasting model to predict the spread of the diseases in the future using time series model, meanwhile the third part is to perform epidemiological modelling and use the result to develop a public health policy to slow down the spread of the infectious disease.
First of all, before getting into the course, we need to ask this question to ourselves: why should we learn biostatistics, particularly infectious diseases analysis? Well, there are many reasons why, firstly, if you are interested in working in the public health or healthcare industry, having biostatistics knowledge would be very beneficial and help you to level up your career. In addition to that, you will also learn a lot of valuable skill sets that can be implemented in other projects, for example, time series decomposition can be used to forecast stock, real estate, commodity, and cryptocurrency markets. Last but not least, this course will also train you to be a better public health policy maker as you will extensively learn how to make data driven decisions and take other external factors into consideration.
Below are things that you can expect to learn from this course:
Learn the basic fundamentals of biostatistics and infectious disease analysis
Learn how to calculate infectious disease transmission rate using SIR model
Learn several factors that accelerate the spread of infectious disease, such as population density, herd immunity, and antigenic variation
Learn how to find and download datasets from Kaggle
Learn how to clean dataset by removing missing rows and duplicate values
Learn how to detect potential outliers using Z score method
Learn how to find correlation between population and disease rate
Learn how to analyze infected patient demographics
Learn how to map infectious disease per county using heatmap
Learn how to analyze infectious disease yearly trend
Learn how to perform confidence interval analysis
Learn how to forecast infectious disease rate using time series decomposition model
Learn how to do epidemiological modeling using SIR model
Learn how to perform public health policy evaluation