
Data is a tool, not a goal.
Every field, every event, and every identifier should exist because it serves a specific segment, trigger, or personalization. If it doesn't, it's an unnecessary cost.
There are three types of data by origin.
Zero-party (intentionally shared by the customer), first-party (captured at the behavioral level), and third-party (from external sources). The strategic foundation of a CDP is the first two.
There are three levels of data by content.
Profile (who the customer is), transactional (how much they're worth), and behavioral (what they do). The minimum starting set is an email or phone number plus a name — full personalization needs all three levels.
Identifiers are the glue of the profile.
Without them, one person turns into five different contacts. The External ID is the most reliable key, since it never changes.
Mistakes cost money.
Duplicates inflate the database without growing the audience. Inconsistent formats break imports. Events without business value add cost without adding value.
Data quality is a matter of discipline, not luck.
The more transparent your collection standards and the more frequent your checks, the cleaner your database — and the higher the quality of personalization and automation.
Data collection is not a one-time task, but an architectural decision.
The chosen collection method determines profile completeness, update speed, the ability to operate in real time, and the system's ability to unify cross-device behavior.
Import and continuous collection are two different processes.
Import defines the depth of the starting point (contacts + events = CDP; contacts only = ESP). Continuous collection is the fuel for triggers, segmentation, and personalization.
Four main collection channels: web tracking, SDK, API, file import.
Each has its own role, strengths, and limitations. In a mature CDP, they work together simultaneously.
Identification is a continuous process.
Each collection method is also an identification method. Cookies, tokens, UTM tags, External ID — all work together so the system understands it is the same person.
Events are the language of a CDP.
Without events, a contact is just a record in the database. An event turns it into a living profile with behavior, history, and automation potential.
Data collection quality = result quality.
Collection errors (missing identifier, incorrect format, broken integration) directly affect triggers, segments, and personalization. Data quality control is an ongoing process.
A list built through dubious means costs more than one built slowly and properly.
A purchased list means legal risk, spam traps, and immediate domain blocklisting. Serious platforms, including Yespo, reject such campaigns, since one spammer harms everyone sharing the same IP subnet.
A legitimate contact is one who has knowingly agreed.
Subscription forms, registration and checkout with a separate consent checkbox, gamified forms, and offline subscriptions tied to documented consent all qualify. Registration or a purchase alone does not constitute consent to receive marketing emails.
Double Opt-In filters out exactly the contacts who shouldn't be on the list.
Addresses belonging to others, bots, and accidental sign-ups get filtered out. Some people won't complete the confirmation, but those who do deliver higher open rates, fewer complaints, and legally documented proof of consent.
Subscription categories turn an all-or-nothing choice into something manageable.
People opt out of the excess rather than opting out of you entirely — fewer unsubscribes, higher open rates, and the category choices themselves become segmentation data.
An honest unsubscribe protects your reputation better than trying to hold people against their will.
A spam complaint does far more damage than an unsubscribe, so the unsubscribe page should offer alternatives — lower frequency or specific categories — rather than a dead end.
Unification is what sets a CDP apart from any other system.
A CRM stores contacts. An ESP sends messages. A DMP works with anonymous audiences. A CDP combines all the data about one customer into a single profile — and that's what creates the value.
Data collection starts before identification.
An anonymous visitor is already a data source. Web tracking captures behavior from the very first visit, and that data becomes part of the profile the moment identification happens.
Merging isn't magic — it's data-transmission logic.
Profile merging happens when both the old (technical) and new (persistent) identifiers are passed in the same request. That's implemented in the integration layer, not inside the CDP itself.
Four rules govern stitching: a single primary ID, normalization, merge control, and source synchronization.
Breaking any one of them leads to duplicates, fragmented profiles, and broken automation.
Unification directly affects segmentation and automation.
An incomplete profile means an inaccurate segment, which means the wrong trigger and irrelevant communication — the chain breaks at the very first link.
Common mistakes are the rule, not the exception.
Broken events, unlinked data, missing historical and offline data, duplicates — every business runs into these. Yespo provides tools to diagnose and fix them: an API for importing events, mechanisms for merging duplicates, and conditional groups for surfacing contacts missing key identifiers.
Data can always be backfilled.
Even if the launch was incomplete, historical events, offline transactions, mobile tokens, and additional attributes can be pushed into the system at any point — a CDP is a system that grows alongside your data.
An event is a signal that a customer's state has changed.
Events are what turn a static contact into a live profile, and they're stored for two years.
System events are automatic; custom events are yours to design.
System events are collected automatically from your campaigns, while custom events come in from outside via API and SDK. There are six categories, and "Other" is the largest, since it holds all behavioral triggers.
The event type key is permanent.
It can't be renamed, only replaced with a new type — so the naming scheme needs to be thought through before the first event, not after the fiftieth.
Every event must carry a contact identifier, or it becomes "orphaned."
The most reliable identifier is the external ID: emails and phone numbers change, the external ID doesn't.
Naming conventions make events self-documenting.
The "object plus action" scheme in the past tense groups events by object. Event names use title case; parameter names use lowerCamelCase — and parameters in workflows must match that casing exactly, since mismatches are the source of countless mysterious failures.
Parameter validation catches integration errors immediately, not a month later.
For anything tied to money, enable it without exception.
Integrations fail silently, so monitoring has to be active.
Events with a new key are automatically registered on the first API request — convenient, but a typo in the key creates a redundant event type rather than an error. Event history, event analytics, and threshold alerts are what surface these issues instead of leaving them hidden.
Segmentation is the point where data turns into action.
Without segmentation, the data you've collected and unified stays as untapped potential. Segmentation turns it into concrete groups with concrete communication strategies.
There are three tiers of segments: strategic, derived, and micro.
Strategic ones are the foundation of the marketing system. Derived ones are for specific campaigns. Microsegments are for triggered automation — each tier has its own role.
There are three levels of depth: basic, advanced, and predictive.
Basic works with static data, advanced with dynamic and behavioral data, predictive forecasts the future. Each level corresponds to a different stage of business maturity.
RFM is the bridge between analytics and automation.
RFM doesn't just show the state of the customer base — in Yespo, it plugs into triggers, so a customer moving between segments automatically fires the relevant communication.
Predictive segmentation answers a different question.
Manual segmentation asks "What did the customer do?" Predictive asks "What will they do next?" — a fundamental shift in how you approach targeting.
Segmentation quality equals data quality.
Everything covered in the earlier lectures — collection, identification, unification — directly affects segment accuracy. An incomplete profile means a flawed segment, which means the wrong communication.
This course contains the use of artificial intelligence.
Most marketing teams collect data long before they know what to do with it — and end up paying for fields, events, and integrations that never drive a single decision. This module fixes that from the ground up.
You'll start with the foundation of any CDP: what data actually deserves to be collected, and why zero-, first-, and third-party data play very different strategic roles. From there, you'll see how that data actually reaches the system — through web tracking, mobile SDK, API, and file import — and how each channel identifies a customer, from an anonymous cookie to a fully recognized profile.
You'll also learn where a contact list should legitimately come from — and why a purchased list is a liability, not a shortcut. This covers consent, Double Opt-In, and subscription categories that keep your list clean and your deliverability intact.
Next, you'll go inside the process that separates a real CDP from a simple contact database: unification. You'll learn how scattered data points — a cookie, a device ID, an email, a phone number — get stitched into one 360-degree customer profile, and why this step directly determines whether your triggers, RFM analysis, and personalization actually work.
From there, you'll look at events — the signals that turn a static profile into a living one. You'll learn how events are structured, named, and validated, and why a poorly designed event scheme quietly breaks triggers and segments down the line.
Finally, you'll turn that unified profile into business action through segmentation — from strategic and derived segments to microsegments, and from basic demographic splits to advanced, RFM, and predictive segmentation, illustrated with real case studies.
By the end of this module, you'll understand data, identity, and segmentation not as separate topics, but as one connected system that determines the accuracy of every campaign built on top of it.