
Discover how data serves as the foundation of data science, explaining what data is and why it matters. Learn about data types, data structures, tabular data, and data life cycle.
Trace the rise of data from ancient tallying to modern data science, cloud computing, and deep learning, highlighting how data storage, processing, and analysis evolved.
Discover how data underpins data science, from data types—categorical and numerical—and representations to tabular data and the data life cycle, guiding the journey from raw data to actionable insight.
Explore what data is and why it matters for data science, and learn how data becomes information, knowledge, and actionable insight for intelligent decision making.
Explore how data are symbols describing observations, including qualities, quantities, and their representation with words and numbers. Learn how organizing, analyzing, and interpreting data turns a datum into information.
Information reduces uncertainty by answering who, what, where, how many, and how much. It is created from data through organizing, analyzing, and interpreting to add context and meaning, becoming knowledge.
Explore how knowledge, the theoretical and practical understanding of the natural world, explains observations, predicts behavior, and guides decisions by combining existing knowledge with new information to solve problems.
Explore how data transforms into actionable insight to support data-driven decisions. Collect, organize, analyze, and interpret data to convert it into information and knowledge for informed action.
Explore data-driven decision making through an investor example: collect apple price data, analyze changes, learn the relation to apple cider prices, and decide to invest to capture profit.
Define data as raw and unorganized facts recorded from observations. Explain information as data organized, analyzed, and interpreted to provide context and knowledge, guiding data-driven decision-making that transforms data into action.
Explore the main data types in data science, including categorical and numerical data, and distinguish nominal, ordinal, interval, and ratio data along with their limitations.
Explore the two main data types in data science, categorical and numerical, and their four subtypes: nominal, ordinal, interval, and ratio data.
Explore nominal data, a type of categorical data with no natural rank order, using examples like colors and names, and apply equality testing and mode to summarize categories.
Explore ordinal data, a categorical data type with a natural rank order. Sort by rank, test equality and order, and determine the median using examples like grades and medals.
Explore interval data, a numerical data type with an arbitrarily chosen zero point, where we can add or subtract but not form ratios, with Celsius, IQ, dates, and longitudes.
Explore ratio data as a powerful type of numerical data with a natural zero point, enabling multiplication, division, and the geometric mean.
Explore the two main data types—categorical and numerical—and differentiate nominal, ordinal, interval, and ratio data. Prepare to learn how these data types are stored in a computer.
Explore data types, binary representations, and how computers store and operate on scalar and composite data. Learn common scalar and composite types you'll encounter in data science.
Explain how computers store data as bits and bytes, and how data types tell computers how to interpret, store, display, and operate on those values, including scalar and composite types.
Explore scalar data types as single-value building blocks of data science, powering storage and operations from letters and numbers to dates, grouped into categorical, numerical, and temporal types.
Explore core scalar data types in data science, including character strings, boolean, enumeration, integer, decimal, float, date, time, date-time, and date-time with time-zone offset.
Explore composite data types in data science as logical containers that group scalar values, providing access methods and operations, organized into homogeneous, tabular, semi-structured, and multi-media data types.
Explore common composite data types in data science, including vectors, matrices, tensors, dictionaries, tables, trees, graphs, and multimedia representations of text, images, audio, and video.
Explore data types and how computers store data, from binary representations to scalar and composite types like characters, integers, dates, times, vectors, tables, graphs, images, and tabular data queries.
Explore tabular data and learn to structure data and extract information from tables using queries, with observations in rows, variables in columns, and cross-table relationships.
Explore tabular data as a table organized in a two-dimensional grid, where each column holds homogeneous data and each row combines heterogeneous values across observations, variables, and relationships.
Identify observations as records in tabular data, with one observation per row, noting that rows may be called tuples or records and include date and measurements like heart rate.
Identify variables as placeholders that vary across observations; store them in columns with consistent data type, scale, and units, keeping one variable per column (date, heart rate, temperature, systolic/diastolic pressure).
Explore how data science represents relationships across tables by using primary keys and foreign keys, keeping cohesive tables for single entities, and linking patient and vital-signs data.
Explore extracting information from tabular data with queries, using SQL, Python, and R, and see a Bill example filtering by name and date in Vital Signs to compute average temperature.
Explore tabular data in rows and columns, focusing on observations and variables. Learn how primary and foreign keys relate tables and how queries reveal information.
Explore the life cycle of data from collection to action. Learn the tools and methods across data collection, storage, processing, analysis, and taking action, guided by feedback-driven optimization.
Observe phenomena, measure quantities, and record observations as binary representations to create data in the data lifecycle; data exist only after observation and recording.
Store data in persistent storage to retrieve it for future analysis. Choose from file-based formats like csv, web-based formats like json, and platforms such as spark and hadoop.
Process data by transforming, cleaning, and querying stored data to prepare reliable analysis, then choose between manual analysis in Excel, scripted workflows, or automated data ETL pipelines.
Analyze data to turn processed information into actionable insights, choosing the right tools—from reports and dashboards to data mining, machine learning, and data-driven artificial intelligence—with Excel, Power BI, and Tableau.
Data analysis drives action within the data lifecycle by informing decisions to act or inaction, implementing changes, and recording outcomes to guide manual and automated decisions.
repeat emphasizes using feedback from outcomes to drive the next iteration, enabling rapid, continuous improvement of business processes through an iterative data science cycle.
Understand the data life cycle from collection to action, including storage, processing, and analysis, and learn how feedback drives iterative optimization.
Wrap up this final module of the data for data science course by practicing through quizzes and exercises, then start the next course and apply what you learned.
See data as the foundation for data science, with data as facts describing observations, and explore categorical and numerical data, data types, tabular data, and the life cycle.
In an information economy, data is the new oil. However, most people lack even a basic understanding of what data are and how they are used in data science. As a result, many of us will be left behind during the next industrial revolution, an information revolution.
In this course, we will learn about data as a foundation for data science. We’ll learn what data are and why they are important. In addition, we’ll learn about data types, data structures, tabular data, and the data life cycle.
By the end of this course, you’ll understand data and how data are used in data science.