
Set up your data science for healthcare coding environment with Anaconda, launch Jupyter notebooks via Anaconda Navigator, and use Python 3 notebooks and markdown for coding and documentation.
Explore Python basics for data science, covering variables, data types, operators, loops, data structures, list comprehensions, and libraries, enabling data wrangling, visualizations, and machine learning.
Discover Python basics by exploring variables, data types, and operators, with examples of patient data and naming conventions, type checks, and arithmetic, comparison, and logical operators.
Explore for loops and while loops to iterate through lists, apply calculations, and illustrate practical examples like dose calculations from patient weights and glucose monitoring alerts.
Explore Python data structures, including lists, dictionaries, sets, and tuples, their properties, ways to access elements with indexing and slicing, and their use in healthcare data.
Master list comprehensions and essential methods, from squaring numbers 0–9 to converting Celsius to Fahrenheit and classifying cholesterol, plus string and list manipulations.
Learn how to create reusable Python functions that encapsulate tasks and simplify coding. Understand def syntax and parameters, and see examples like add numbers, greet patients, and a BMI calculator.
Learn object oriented programming by building a patient class with a constructor, self, attributes, and methods to manage allergies and GFR in Python.
Acquire and import data from files, APIs, web scraping, and databases like SQLite and Postgres, using pandas to load csv, excel, json into a unified data frame.
Acquire and import data from CSV, Excel, JSON, and HTML sources with pandas, merge data across sources, inspect datasets with head, and fetch API data via constructed URLs.
Learn to acquire and import healthcare data using web scraping with BeautifulSoup, extract HTML tables into pandas data frames, and query data from SQLite and Postgres databases.
Explore data wrangling in module four by cleaning messy data, handling missing values, removing duplicates, converting data types, and performing a deep copy to safeguard data for machine learning.
Learn data wrangling in healthcare by creating a deep copy, handling missing data, performing type conversion, and removing duplicates with pandas to preserve the original dataset.
Learn to perform type conversion and rounding with pandas, remove duplicates, validate clean data, and export a csv with index disabled for reliable healthcare analytics.
Explore exploratory data analysis in module five using pandas to generate descriptive statistics, visualize data with histograms, box plots, and scatter plots, and assess outliers with domain knowledge.
Explore healthcare data with descriptive statistics and visualizations—histograms, box plots, scatter and pair plots—and outlier detection to reveal relationships before modeling.
Identify and remove outliers in health data using the interquartile range in exploratory data analysis, focusing on blood pressure, heart rate, and hemoglobin A1c, then prepare cleaned data for pre-processing.
Explore data pre-processing for healthcare data science, including feature selection, min-max scaling, feature engineering such as body mass index, and one-hot encoding for machine learning.
Learn to preprocess healthcare data for machine learning by creating a risk category target, applying feature scaling with min-max normalization, and preparing a clean dataset for models like random forest.
Standardization using z-score scaling centers data at zero with unit variance, while feature engineering creates age groups, applies one-hot encoding, and drops unused columns to ready data for modeling.
Explore a brief overview of machine learning in healthcare data science. Learn about random forest models, train-test splits, model training, predictions, and evaluation.
Explore how machine learning uses cleaned healthcare data to learn patterns, build models like random forest, and predict outcomes, diagnose diseases, and personalize treatment while considering data quality and privacy.
Apply a random forest classifier to patient data to predict high risk versus other risk levels, using features like blood pressure and heart rate.
Evaluate random forest model performance with precision, recall, F1, and accuracy, focusing on high risk detection, and prepare results and visualizations for module eight.
Master data visualization techniques and data types, apply best practices for visualization, and sharpen presentation skills with tips and an optional final project plus personalized feedback.
Explore best practices for visualizing healthcare data across data types, using appropriate charts, consistent color schemes, and accessible formatting to create clear, narrative presentations.
This course is designed to introduce healthcare students and professionals to the fundamentals of data science using Python. It covers a wide range of topics from basic programming skills to advanced data analysis and machine learning techniques, all within the context of healthcare applications. Through interactive Jupyter notebooks, participants will learn how to analyze and interpret complex healthcare data, ultimately gaining insights that can inform clinical decisions, enhance patient care, and drive healthcare innovation.
Target Audience
• Healthcare students (medicine, pharmacy, nursing, public health, etc.)
• Healthcare professionals (clinicians, researchers, administrators) looking to
incorporate data science into their practice
Prerequisites
• Basic understanding of statistics
• Basic computer skills
• No prior programming experience required
By the end of this course, participants will be able to:
• Set up and navigate the Jupyter notebook environment.
• Understand and apply Python basics for data science.
• Acquire and import healthcare data from various sources.
• Clean and preprocess data for analysis.
• Perform exploratory data analysis (EDA) to uncover trends and patterns.
• Apply data preprocessing techniques to prepare data for modeling.
• Build and evaluate a basic machine learning model to predict health risk.
Course Structure
The course contains eight modules in total, covering topics from an introduction and setting up your coding environment to building and evaluating a machine learning model.
Each module includes example coding notebooks, videos, quizzes, and supplementary notes to enhance your learning experience.