Identity Graph
Many identifiers, one entity - the map that makes cross-device targeting, suppression, and measurement possible, and probabilistic where it isn't careful.
- Term
- Identity Graph
- Maps
- Emails, devices, cookies → one entity
- Built by
- Deterministic + probabilistic matching
- Privacy line
- Pseudonymous, consented, vendor-diligenced
Forms & parts of speech
Definition in plain terms
An identity graph is a database that maps the many identifiers belonging to one person or household — email addresses, device IDs, cookies, phone numbers, account IDs, the HASHED-EMAIL join keys — into a single connected entity. It is the engine under cross-device targeting, frequency capping, suppression, personalization, and measurement: the thing that knows the phone, the laptop, and the logged-out cookie are the same human (the IDENTITY-RESOLUTION process is how the graph gets built and maintained).
The mechanics
The build splits the way the matching entries do: DETERMINISTIC edges (joins on shared, verified identifiers — the same hashed email or login across systems — high-confidence, the spine where it exists), and PROBABILISTIC edges (statistical inference that two devices are the same person from shared IP, behavior, and timing — broader coverage, lower confidence, accuracy that varies by vendor and degrades silently). Identity graphs come in flavors that decide their trustworthiness: first-party graphs (your own — the GOLDEN-RECORD and SINGLE-CUSTOMER-VIEW built from owned, consented data, the most accurate and defensible), and third-party graphs (vendor-built across the ecosystem — broad reach, opaque methods, and the supply-chain diligence burden the FTC's data-broker enforcement made non-optional). The accuracy truth that governs everything downstream: a graph is probabilistic infrastructure whose error rate becomes every connected system's error rate — wrong edges merge two people (the suppression that leaks, the personalization that creeps, the frequency cap that punishes the wrong household) and missing edges fragment one person (the duplicate-send and split-LTV problems the golden-record entry details) — so the graph deserves the accuracy audit nobody schedules. The privacy frame this entry must carry: identity graphs are pseudonymous-personal-data systems under GDPR-grade law (linking identifiers IS the processing the law governs), consent must travel into and through them, the HOUSEHOLD-level versions are sensitive, and the post-cookie rebuild raised their strategic stakes — as THIRD-PARTY-COOKIE and IDFA deterministic signals decayed, first-party identity graphs became the durable identity asset, which is exactly why their accuracy and governance stopped being a vendor footnote.
When it matters
Identity graphs matter wherever marketing acts on 'the same person' across devices, sessions, or channels — cross-device targeting and frequency, suppression and consent enforcement, personalization, clean-room matching, and unified measurement — which the cookie's decline made the central infrastructure problem rather than a plumbing detail. They matter most as an accuracy-and-governance responsibility (the graph's error and consent posture propagate everywhere) and as a build-versus-buy decision (first-party defensibility versus third-party reach). The discipline is deterministic spine where possible, probabilistic edges priced for their accuracy, third-party graphs diligenced like the broker data they are, consent carried through, and the audit run before the wrong edge becomes a privacy incident.
Synonyms & antonyms
Synonyms
Antonyms
Origin & history
Identity graphs grew from CRM matching and the ad-tech cookie-syncing era into the post-cookie identity stack's core asset; the social platforms proved the deterministic logged-in version at scale, and the THIRD-PARTY-COOKIE and IDFA decline turned first-party identity graphs from a vendor luxury into the durable spine of addressable marketing.
Etymology: source.
Usage trends
Search interest for this term over the last five years:
Common questions
- What is an identity graph?
- A database mapping the many identifiers of one person or household — emails, devices, cookies, accounts — into a connected entity; the engine under cross-device targeting, frequency, suppression, and measurement.
- How are identity graphs built?
- From deterministic edges (verified shared identifiers like a hashed email, high-confidence) and probabilistic edges (statistical inference from IP, behavior, timing — broader but lower-confidence and vendor-variable).
- What is the main risk?
- Accuracy propagates — wrong edges merge two people (leaking suppression and personalization), missing edges fragment one; plus the privacy burden, since graphs are pseudonymous personal data requiring consent and vendor diligence.
Related tools & calculators
- toolCAC calculator
- toolLTV:CAC calculator
Resources & people to follow
- referenceIAB — identity and addressability resources
- referenceFTC data-broker and identity enforcement actions
- referenceRGM analysis — the graph's error rate is every downstream system's error rate; audit it, govern it, prefer first-party
Curated, non-competitor resources verified per term.
Related training
- modulePerformance marketing
Disciplines
Areas of marketing where identity graph is a core concern: