Skip to content
All Resources
Identity Data 12 min read

Prepared by GSDSI Regulatory Content Team

Last updated

Editorial standardsTrust Center

What Deterministic vs Probabilistic Matching Costs

Deterministic matching is a join on the same identifier after agreed normalization. Probabilistic matching infers a link from other evidence. Mixing them in one match rate hides false joins, missed suppression, and inflated reach. Require relationship-type labels, score each lane on the same seed, and put the mix rule in the contract.

How to use this article

Read the checklist here, then use the linked hub and product pages for procurement citations.

What the Labels Actually Mean

Use operational definitions, not brand names. Deterministic means both sides present the same identifier after a documented transformation: the same email after lowercasing, trimming, and hashing; the same MAID; the same customer ID; the same VIN. The join does not guess who the person is. It tests equality of keys. Probabilistic means the vendor infers a link from other evidence: name and postal similarity, device co-occurrence, household graphs, or a score. That inference can be useful. It is still an inference. If a vendor calls a scored name/address link deterministic, they have renamed the method. Put the method in the exhibit.

Identity products often contain both. MAID identity joins and Core Email File appends can be deterministic on the key you send and probabilistic on the extra endpoints the graph returns. Identity graph pages should be read as a map of relationship types, not as a promise that every edge is an equality join. For MAID-to-HEM, hashing is a join and transport control on that product. It does not convert a household inference into a deterministic person match. Raw and hashed data are otherwise licensed under written terms in the broader catalog.

IAB Tech Lab identifier work is a useful external vocabulary for key types. NIST Privacy Framework is useful when you need to describe disassociability and inventory for a scored link that expands a seed into additional people or devices. Neither standard tells you what a specific vendor's score means. That documentation has to come from the vendor, per SKU, with a way to turn the score off.

What False Joins Cost You

A false join is a link that should not have been made for your use. In audience targeting, it looks like reach you did not intend: ads to the wrong household member, the wrong device, or a neighbor who shared a last name and a building. In measurement, it looks like inflated incrementality because exposure and outcome were attached through a weak edge. In suppression, it looks like a deletion or opt-out that did not propagate to the extra endpoints the graph invented. Those are operational costs. They show up in wasted media, noisy lift, and complaints. They do not show up in a blended match rate, because the false join still counts as a match.

Where false joins land in a buying workflow
LaneWhat breaksWhat to measure in the pilot
ActivationReach and frequency on the wrong endpointsCapped vs uncapped household expansion; sampled false-link review
MeasurementOutcomes attached through weak edgesReport with probabilistic links on and off
SuppressionOpt-out or deletion misses expanded IDsSeed a known suppression key and inspect the output
EnrichmentWrong email or phone on a true customerHoldout of known-good contacts; conflict rate

Household expansion deserves its own row because it is often sold as deterministic when the household grouping is stable and the use is not. A TV ID that correctly belongs to a household is still the wrong grain if your suppression list is person-level. A MAID that co-locates at an address is still a weak person match if the campaign is meant to reach one named customer. The CTV attribution last-mile guide treats this as a measurement problem. It is also a matching-cost problem: you pay for extra endpoints and you inherit extra risk. Require the vendor to label household edges and to run the pilot with expansion off.

Do not debug false joins only with anecdotes. Pre-register a messy cohort: shared last names, multi-unit addresses, common business names, recycled MAIDs, and role-based emails. Score precision on that cohort separately from the clean cohort. The seed match testing guide is the sample-design companion. If the vendor's probabilistic lane looks strong only on the clean cohort, you learned something checkable.

What Missed Joins Cost You

A missed join is a true relationship the vendor did not return. The cost is the opposite of false-join waste: you understate overlap, you re-buy the same people through another channel, and you conclude the vendor is thin when the failure was prep or recency. Common causes are undocumented hash rules, stale keys, destination rejection, and a deterministic-only run on a seed that needed a documented probabilistic assist you never asked for. The fix is still a split report, not a higher blended rate.

If you already have a first-party spine, missed joins are easy to see. Send known-good keys. If the deterministic lane misses them, the issue is prep, vintage, or coverage, and you can ask for evidence on each. If the deterministic lane hits them and the probabilistic lane adds a crowd of extra endpoints, you are looking at expansion, not recall. Keep those numbers apart. The MAID graph diligence guide and why vendor match rates are not comparable cover denominator tricks that make missed joins look like high coverage.

Measurement teams feel missed joins as broken incrementality in the other direction: true exposed households that never enter the matched conversion set. MRC standards are a useful external reference for what a measurement claim is supposed to mean. They will not tell you whether a vendor's scored edge belongs in that claim. Your pilot should produce two measurement files, one with inferred links excluded, so finance can see the sensitivity. The identity graph activation playbook is the adjacent workflow for email-to-CTV paths where this choice is explicit.

How to Test Both Lanes on One Seed

You do not need two seeds. You need two labeled outputs from one seed, plus a holdout the vendor cannot tune. Ask for a deterministic-only file and a file that includes probabilistic or household links, with a relationship-type and confidence field on every extra edge. If the vendor cannot emit that field, they cannot support the contract language you need. Run the same spec against MAID and email separately so a strong lane in one product cannot hide a weak lane in the other.

  1. Write the deterministic rule: which keys, which normalization, which hash.
  2. Write the probabilistic rule: which evidence, which score threshold, how to disable it.
  3. Write the household rule: cap, grain, and whether it is on by default.
  4. Send one seed. Require two (or three) output files with row counts that reconcile.
  5. Include a messy cohort and a holdout. Review a sample of extra edges, not only the rate.
  6. Seed a suppression key. Confirm it lands on every endpoint in every lane.
  7. Re-run later in the evaluation. Inferred graphs drift. Equality joins should not, unless the spine changed.

Price talks should follow the same split. If probabilistic endpoints are cheaper per thousand, ask whether they are cheaper because they are inferred. Paying less for a weaker edge can still be more expensive in media waste. The enterprise pilot checklist is where delivery format, cadence, and production drift belong. Do not let a blended CPM collapse the lanes you just separated in QA.

Contract Language That Keeps the Lanes Separate

If the method is not in the order form, production will ship the vendor's default mix. Require: relationship-type on every delivered link; a way to order deterministic-only; a default-off rule for household expansion unless the use is household-grain; suppression propagation to expanded endpoints; and a retest right on the same seed after a methodology change. The licensing red-flag guide covers derived data and audit rights that make those clauses enforceable.

Permitted use follows the grain. A person-level consent or suppression flag does not automatically authorize every device and household endpoint the graph can emit. Spell the expansion rule in the same exhibit as permitted use. For activation programs, keep audience targeting destinations named, because destination acceptance is part of the match. For measurement programs, name whether inferred links may enter the outcome join. Questions on a scoped identity sample: contact.

Frequently Asked Questions

Is a hashed-email join deterministic?
It is deterministic if both sides present the same email after the agreed normalization and hash, and the join is equality on that key. A scored name and postal link that also happens to return an email is still probabilistic. Hashing does not change the method.
Should household expansion be treated as deterministic?
Not for a person-level use. A stable household grouping can still attach extra people and devices. Report household edges as their own lane and run the pilot with expansion off if the use is person-grain.
How do false joins show up in suppression?
A deletion or opt-out on the seed key may not flow to inferred endpoints unless the vendor labels those edges and the contract requires propagation. Test with a known suppression key.
Can I price deterministic and probabilistic links the same way?
You can, but you will hide cost. Inferred endpoints can look cheaper per thousand and still cost more in wasted media and noisy measurement. Keep the lanes in separate columns of the scorecard.
What is the minimum identity-graph pilot output?
One seed, a deterministic-only file, a file that includes inferred or household links with relationship-type labels, a suppression test, a holdout review of extra edges, and a re-run during the evaluation.

✓ Opt-Out Request Honored via Global Privacy Control