Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
IBM SPSS Modeler: Techniques for Missing Data
Rating: 3.8 out of 5(27 ratings)
278 students

IBM SPSS Modeler: Techniques for Missing Data

IBM SPSS Modeler Seminar Series
Created bySandy Midili
Last updated 4/2014
English
English [Auto],

What you'll learn

  • Understand how missing data is identified and defined in IBM SPSS Modeler
  • Impute missing values
  • Remove missing data
  • Run parallel streams with and without missing data
  • Use the Type, Data Audit, Derive, and Filler nodes to identify and handle missing data

Course content

2 sections20 lectures3h 16m total length
  • Introduction to Missing Data4:21

    Learn definitions and evaluation of missing data and how to address it with IBM SPSS Modeler by using Modeler nodes and practical imputation tips.

  • Missing Data within the context of CRISP-DM5:27

    Assess data quality in the data preparation stage by examining distributions, outliers, and how much missing data you have, and decide how to handle missing values to improve model accuracy.

  • Reasons for Missing Information3:44

    Identify why data may be missing, including privacy concerns, sensitive questions, memory lapses, or data entry errors, and remember that some fields may not apply or be lost.

  • Type and Amount of Missing Data6:33

    Assess the type and amount of missing data to guide handling decisions. Identify patterns and reasons behind gaps, then apply options based on missing data type and impact on analysis.

  • Missing Data Issues6:19

    Explore missing data issues that affect both sample size and data quality, including nonresponse bias, the importance of key fields, and population considerations for generalizable models.

  • Ways to Address Missing Data10:13
  • Missing Data Definitions2:18

    Explore how IBM SPSS Modeler defines missing data, including null values for numeric fields, whitespace for categorical fields, and predefined blank codes like 99.

  • Useful Nodes to Handle Missing Values14:48

    Define missing values with the type node, then impute or remove fields and cases, and use the data audit to inspect distributions and completeness.

  • A First Look at the Data9:47

    Explore how to identify, define, and visualize missing data in IBM SPSS Modeler, audit data quality, and apply fixed, random, and algorithmic imputations on a real dataset.

  • Removing Fields and Records27:09

    Explore methods to handle missing data by removing fields or records, imputing values, and deriving flags, including using percent complete, filter nodes, and total missing counts to guide data preparation.

  • Creating Null Flags11:56

    Create null flags for each variable to test whether missingness predicts the target, identify top missingness indicators, and prioritize variables for imputation and feature selection.

  • Imputing with the Data Audit Node24:47

    Explore imputing missing data with the data audit node in IBM SPSS Modeler, using fixed values, mean, random and predictive imputation, and coercion to preserve distributions.

  • Using Full and Partial Data8:05

    Learn to handle missing data by using full and partial data with a two-model approach that separates new donors from non-new donors, aided by feature selection.

  • Imputing the Median and the Mean10:51

    Impute the median and the mean using aggregate and merge steps in IBM SPSS Modeler to replace missing age values, compare distributions, and explore normal distribution imputations.

  • Using the Anti-Join11:12

    Explore how to merge data using joins, review inner and left join options, handle missing records, and assess the impact of dropping rows by examining NDA variables and revenue implications.

Requirements

  • Knowledge or experience with IBM SPSS Modeler or completion of an introductory level data mining course and on the job data mining experience.

Description

IBM SPSS Modeler is a data mining workbench that allows you to build predictive models quickly and intuitively without programming. Analysts typically use SPSS Modeler to analyze data by mining historical data and then deploying models to generate predictions for recent (or even real-time) data.

Overview: Techniques for Missing Data is a series of self-paced videos (three hours of content). Students will learn how missing data is identified and handled in Modeler. Students also will learn different approaches to dealing with missing data including imputation of missing values, removing missing data, and running parallel streams with and without missing data. Students will also learn how to use the Type, Data Audit, and Filler nodes to identify and handle missing data.

Who this course is for:

  • Anyone that has experience with IBM SPSS Modeler or has completed an introductory level data mining course and would like to learn about different ways to handle missing data.