Growth Marketing Glossary

Identity Graph

i·den·ti·ty graphnoun

Many identifiers, one entity - the map that makes cross-device targeting, suppression, and measurement possible, and probabilistic where it isn't careful.

emailcookiephonedeviceIDonepersonidentifiers stitchedto an entitythe map linking scattered identifiers to one person or household
Schematic — identifiers stitched to one entity
Term
Identity Graph
Maps
Emails, devices, cookies → one entity
Built by
Deterministic + probabilistic matching
Privacy line
Pseudonymous, consented, vendor-diligenced

Forms & parts of speech

identity graph · noun
The identifier-stitching map.
"The identity graph decided everything downstream - match rates, frequency, suppression - and nobody had audited its accuracy in two years."

Definition in plain terms

An identity graph is a database that maps the many identifiers belonging to one person or household — email addresses, device IDs, cookies, phone numbers, account IDs, the HASHED-EMAIL join keys — into a single connected entity. It is the engine under cross-device targeting, frequency capping, suppression, personalization, and measurement: the thing that knows the phone, the laptop, and the logged-out cookie are the same human (the IDENTITY-RESOLUTION process is how the graph gets built and maintained).

The mechanics

The build splits the way the matching entries do: DETERMINISTIC edges (joins on shared, verified identifiers — the same hashed email or login across systems — high-confidence, the spine where it exists), and PROBABILISTIC edges (statistical inference that two devices are the same person from shared IP, behavior, and timing — broader coverage, lower confidence, accuracy that varies by vendor and degrades silently). Identity graphs come in flavors that decide their trustworthiness: first-party graphs (your own — the GOLDEN-RECORD and SINGLE-CUSTOMER-VIEW built from owned, consented data, the most accurate and defensible), and third-party graphs (vendor-built across the ecosystem — broad reach, opaque methods, and the supply-chain diligence burden the FTC's data-broker enforcement made non-optional). The accuracy truth that governs everything downstream: a graph is probabilistic infrastructure whose error rate becomes every connected system's error rate — wrong edges merge two people (the suppression that leaks, the personalization that creeps, the frequency cap that punishes the wrong household) and missing edges fragment one person (the duplicate-send and split-LTV problems the golden-record entry details) — so the graph deserves the accuracy audit nobody schedules. The privacy frame this entry must carry: identity graphs are pseudonymous-personal-data systems under GDPR-grade law (linking identifiers IS the processing the law governs), consent must travel into and through them, the HOUSEHOLD-level versions are sensitive, and the post-cookie rebuild raised their strategic stakes — as THIRD-PARTY-COOKIE and IDFA deterministic signals decayed, first-party identity graphs became the durable identity asset, which is exactly why their accuracy and governance stopped being a vendor footnote.

When it matters

Identity graphs matter wherever marketing acts on 'the same person' across devices, sessions, or channels — cross-device targeting and frequency, suppression and consent enforcement, personalization, clean-room matching, and unified measurement — which the cookie's decline made the central infrastructure problem rather than a plumbing detail. They matter most as an accuracy-and-governance responsibility (the graph's error and consent posture propagate everywhere) and as a build-versus-buy decision (first-party defensibility versus third-party reach). The discipline is deterministic spine where possible, probabilistic edges priced for their accuracy, third-party graphs diligenced like the broker data they are, consent carried through, and the audit run before the wrong edge becomes a privacy incident.

Worked example. A retailer's martech runs on a third-party identity graph nobody has audited, and the cracks surface as separate tickets that share one root: a customer gets a 'we miss you' win-back while actively shopping (a missing edge split her devices into two people), a sensitive-category recommendation follows a shared household tablet to the wrong family member (a wrong probabilistic edge merged two people), and clean-room match rates are mediocre with no one able to say why. The fix starts by reading the graph as the load-bearing infrastructure it quietly became: an accuracy audit against a known-truth sample exposes a probabilistic error rate well above what the vendor implied, so the architecture shifts toward a first-party graph - deterministic spine on consented hashed-email logins, probabilistic edges kept but confidence-scored and used only where a wrong join is cheap, the third-party vendor demoted to reach extension with diligence attached, and consent wired to travel through the graph rather than around it. Suppression starts binding to real people, frequency caps to real households, personalization stops creeping, and match rates rise on the deterministic core. The graph was always deciding everything downstream; the audit just made the company start treating it that way.
Failure modes to watch. Probabilistic graphs trusted at deterministic precision until wrong edges leak suppression and personalization; missing edges fragmenting one customer into many (duplicate sends, split LTV); third-party graphs onboarded without the broker diligence regulators demand; consent imagined to stop at the graph's edge rather than travel through it; and the accuracy audit nobody schedules on the system every downstream metric depends on.

Synonyms & antonyms

Synonyms

identity graphidentity mapcross-device graph

Antonyms

siloed identifiers (no graph)deterministic-only graph (precise, narrow)

Origin & history

Identity graphs grew from CRM matching and the ad-tech cookie-syncing era into the post-cookie identity stack's core asset; the social platforms proved the deterministic logged-in version at scale, and the THIRD-PARTY-COOKIE and IDFA decline turned first-party identity graphs from a vendor luxury into the durable spine of addressable marketing.

Etymology: source.

Usage trends

Search interest for this term over the last five years:

View interest-over-time on Google Trends →

Common questions

What is an identity graph?
A database mapping the many identifiers of one person or household — emails, devices, cookies, accounts — into a connected entity; the engine under cross-device targeting, frequency, suppression, and measurement.
How are identity graphs built?
From deterministic edges (verified shared identifiers like a hashed email, high-confidence) and probabilistic edges (statistical inference from IP, behavior, timing — broader but lower-confidence and vendor-variable).
What is the main risk?
Accuracy propagates — wrong edges merge two people (leaking suppression and personalization), missing edges fragment one; plus the privacy burden, since graphs are pseudonymous personal data requiring consent and vendor diligence.

Related tools & calculators

Resources & people to follow

Curated, non-competitor resources verified per term.

Related training

Disciplines

Areas of marketing where identity graph is a core concern:

Sources

  1. trendsGoogle Trends — "identity graph"