Choosing an identity graph comes down to five checks: which identifier types it resolves, how much of it is deterministic versus probabilistic, how far its coverage reaches domestically and internationally, whether delivery and refresh fit your stack, and what each identifier key is licensed for. Match rate matters, but only after these five line up.
How to use this article
Read the checklist here, then use the linked hub and product pages for procurement citations.
Teams shopping for an identity graph usually open with a match rate and a record count. Both are fair questions, and both come too early. A graph can post a strong number and still be the wrong product, because it links the wrong keys, reaches the wrong markets, or arrives in a shape your warehouse cannot use. Work through the list below first, then argue about match rates.
1. Which Identifier Types Does the Graph Resolve?
"Identity graph" covers very different products. Some link mobile advertising IDs (MAIDs) to hashed emails (HEMs). Some group devices into households. Some stitch cross-device activity to a single person. Some do all three, at different levels of confidence. Write down the pairing your use case actually needs before you look at a single vendor deck.
If you are activating mobile audiences from a CRM file, you need device-to-email linkage. If you are measuring household reach on connected TV, you need household grouping. A graph that is excellent at one and silent on the other is a mismatch, not a bargain. The test is simple: list your input keys and your output keys, and ask the vendor to show which of those pairs they resolve directly.
Put that pairing in writing. A one-page compatibility matrix, your input keys down one side and the graph's resolvable outputs across the top, turns a sales conversation into a specification your data engineering team can hold the vendor to after signature, not just during the pitch.
2. Deterministic, Probabilistic, or a Blend?
Deterministic links come from a shared, observed identifier. Probabilistic links are inferred from signals such as IP address, timing, and behavior. Neither is wrong. They are tools with different failure modes, and the blend a vendor uses tells you what kind of errors you will be paying to absorb. A graph that is mostly inferred tends to reach further and mislink more.
Ask for the mix per identifier pair, not a single blended label. Then decide where your use case can tolerate inference: broad prospecting usually can, suppression and measurement usually cannot. The operational cost of deterministic versus probabilistic matching is worth reading before this conversation. How to turn that mix into enforceable match rate and decay terms is its own topic, covered in our guide to identity graph SLAs.
3. Where Does the Coverage Actually Reach?
A single headline number hides where a graph is deep and where it is thin. Ask for coverage split by country and, for US buyers, by market. Domestic and international reach are usually built from different supply and should be quoted separately. For reference, GSDSI's MAID Feed catalogs roughly 200M+ US MAID-to-HEM linkages and 700M+ international device IDs across 150+ countries, listed as two figures because they describe two different populations.
Then test the slice you care about. If most of your spend sits in a handful of markets, national coverage is trivia. A seed-match test on your own records in those markets tells you more than any catalog page.
4. Will It Land Cleanly in Your Stack?
An identity graph you cannot load is shelfware. Confirm the file formats and destinations before procurement signs: flat files, columnar formats, cloud storage drops, SFTP, or a warehouse share. The MAID Feed, for example, ships as CSV or Parquet through S3, SFTP, or a Snowflake share. Ask how schema changes are announced, too, because a renamed column breaks a pipeline faster than a coverage gap.
Refresh cadence belongs on the same list. Identifiers reset and devices change hands, so a graph loaded once and left alone drifts away from reality. A daily refresh, which is how the MAID Feed is published, only helps if your ingestion runs on a matching schedule. Decide whether you want full replacements or incremental deltas, and size your jobs accordingly. This is what separates a graph that powers audience targeting and cross-channel measurement from one that sits in a bucket.
5. What Is Each Identifier Key Licensed For?
Permitted use is not a property of the graph. It is a property of each key and its source. A HEM collected under one consent flow may be cleared for activation, while a MAID from a different supply path is cleared only for measurement. Ask for a permitted-use matrix by identifier type, and make sure your legal team reads it before your engineers build on it.
Run these five checks in order and the shortlist usually shrinks to the few graphs that fit your keys, your markets, and your stack. That is the point to compare match rates on your own data. If device-to-person linkage is the gap you are filling, request a MAID Feed sample and run it against your file.
Frequently Asked Questions
What should I check first when choosing an identity graph?
Start with the identifier types it resolves. List the keys you have and the keys you need, and confirm the graph links those exact pairs, whether that is MAID to HEM, device to household, or cross-device to person. Coverage and match rate only matter once the pairing fits.
Is a probabilistic identity graph worse than a deterministic one?
Not automatically. Probabilistic links extend reach but carry more mislinks, while deterministic links are narrower and more reliable. The right mix depends on the job: prospecting can tolerate inference, while suppression and measurement usually need deterministic links.
Why should domestic and international coverage be quoted separately?
They are usually built from different supply and describe different populations. A blended figure can hide a deep domestic graph with thin international reach, or the reverse. Ask for country-level coverage and test the markets where your spend actually sits.
Does refresh cadence matter if the match rate is high?
Yes. Identifiers reset and devices change owners, so a graph that is not refreshed drifts out of date. A daily refresh only pays off if your own ingestion keeps pace, so align the vendor's cadence with your pipeline schedule.
How is this different from a MAID feed diligence checklist?
This checklist is for choosing the identity graph category itself: the identifier types it resolves, its deterministic versus probabilistic mix, and where its coverage reaches. MAID feed diligence, covered in our five-question guide, goes a level deeper on source-app transparency and consent architecture once you have already decided a MAID-based graph fits.
✓ Opt-Out Request Honored via Global Privacy Control
We use cookies and similar technologies to improve your experience. No marketing cookies are pre-selected. Learn more in our Privacy Policy. Your Privacy Choices.