A vendor match rate is a ratio, not a fact about the market. You cannot compare two rates until both vendors use the same definition of a match, the same denominator, the same recency window, and the same rule for household expansion. Write those four down, then run one seed file through every finalist.
How to use this article
Read the checklist here, then use the linked hub and product pages for procurement citations.
Procurement decks treat match rate as if it were a shared unit, like a currency. It is not. The only checkable form is: matches divided by a stated denominator, where both sides are defined. If a vendor will not write that sentence, you do not have a rate. You have a marketing label. Teams licensing MAID identity, Core Email File, or audience targeting should put the formula in the RFP response exhibit, not in a slide footnote.
Start with the numerator. Does a match mean the vendor returned any identifier, a specific identifier you named, a deterministic key after agreed normalization, or a scored link above a threshold the vendor chose? Those are different events. A household graph match is not a person match. A device match is not an email match. A match that later fails destination onboarding is not a usable match for activation. IAB Tech Lab standards are a useful vocabulary for identifier types. They do not replace a product-specific definition in your contract.
Match-rate formula components to pin before a bake-off
Component
Question to ask
Fail signal
Match definition
What identifier, at what relationship level, counts as a hit?
Any returned row counts as a match
Denominator
Input rows, valid keys, opted-in keys, or destination-accepted keys?
Vendor will not name the base
Recency window
As-of date of the file and how stale keys are treated
No vintage on the match report
Household rule
Are person, device, and household links reported separately?
One blended rate for all three
Prep spec
Hash, case, trim, and encoding rules for the seed
Vendor prep is undocumented
Hashing and normalization change the numerator before any graph is involved. If one vendor lowercases and trims email, and another does not, you are not measuring coverage. You are measuring prep. Document the transformation, keep a copy of the exact seed bytes you sent, and require the vendor to echo the prep rules in the match report. For MAID-to-HEM joins, hashing is a join and transport control on that product. It is not a claim that every GSDSI product is hashed-only. Raw and hashed data are both licensed in the broader catalog under written terms.
Why Denominators Diverge Across Vendors
The same seed can produce different rates because vendors drop different rows before they start counting. Common bases include: all input rows; rows that passed format validation; rows the vendor classified as in-universe; rows with a live advertising identifier; rows that survived suppression; and rows a destination later accepted. Each step shrinks the denominator and lifts the reported rate if the vendor starts counting after the drop. Ask where counting begins. Then ask for the count at every prior step so you can rebuild the ratio yourself.
A useful report therefore has more than one rate. At minimum, separate input validity, eligible resolution, destination acceptance, household expansion, suppression, and unmatched. The MAID identity-graph diligence guide walks those stages for device and hashed-email graphs. The same split applies to email append, B2B contact match, and CTV onboarding. If a vendor insists on one blended percentage, treat that as a documentation failure, not as a stronger product.
Destination acceptance is easy to miss. A graph can resolve a seed to a device or household identifier that the activation platform will not take. That is not a match for the workflow you are buying. If the use is audience targeting into a named DSP or clean room, the match report should include the destination's accepted-key count, not only the vendor's internal join count. The enterprise pilot checklist is the place to lock that destination before the sample runs.
How to Read a Coverage Claim Critically
Coverage is the cousin of match rate, and it fails in the same way: the universe is unnamed. Coverage of what? Adults, households, smartphones with advertising IDs, licensed drivers, businesses with a named contact, or records that exist in a file regardless of whether they are reachable? A claim that does not name the universe cannot be checked. A claim that names the universe but not the as-of date cannot be checked against decay. A claim that includes modeled or expanded links without labeling them cannot be checked against a deterministic spine.
Universe: who or what is in the claimed set, in operational language, not a brand name.
Vintage: as-of date of the extract you are evaluating, not the date of the sales deck.
Observation vs model: which links were seen, which were inferred, and whether they are separable.
Reachability: whether the record is joinable, contactable, or only present as an archive row.
Overlap: whether the vendor is counting unique keys or household-expanded endpoints.
Identity-graph coverage is a frequent place this goes wrong. A graph can be large and still miss the cohort you actually activate: recent movers, iOS-reset devices, B2B domains, or a geography your campaign needs. Headline graph size is not a substitute for a seed test on that cohort. The identity graph overview describes join types. Your diligence packet should still require a cohort split on the seed you care about, not a national average the vendor prefers. NIST Privacy Framework language is useful when you need to connect those coverage limits to inventory, control, and disassociability questions for security review.
Do not accept a coverage map that is only a choropleth of where the vendor has any records. Ask for the same map restricted to keys that survived your match definition and recency window. If the vendor cannot produce that view, you cannot tell whether the colorful regions are live inventory or historical residue. For clickstream and intent overlays, pair the identity test with clickstream web intent recency rules so a strong identity match is not carrying a stale behavior flag.
What a Matched-Sample Report Should Split
A matched sample is not a demo file the vendor selected. It is your seed, prepared to a written spec, run through the candidate products, with a report that reconstructs the funnel. The seed match testing guide covers sample design, messy cohorts, and holdouts. This section is only about the report shape that makes two vendors comparable.
Echo the seed row count and the prep rules the vendor applied.
Report valid-key count after format and encoding checks.
Report eligible resolution by relationship type: person, device, household, or account.
Report destination-accepted count for the named activation or measurement path.
Report suppression and deletion hits as their own line, not as unmatched.
Report unmatched with a reason code where the vendor has one.
Re-run the same seed later in the evaluation and explain any material change.
If you are scoring MAID and email in the same bake-off, do not blend them. A strong email append and a weak device join are different products. Keep Core Email File results in their own column. If household expansion is in scope, produce a second table with expansion capped and uncapped so you can see how much of the rate is extra endpoints rather than extra people. The RFP scorecard should weight those columns separately. A vendor that will only return a single blended rate is asking you to fund opacity.
Questions to Paste Into the RFP
These questions are checkable because a complete answer is a document you can replay on a seed, not a slogan. Paste them into the RFP as a required exhibit. Reject oral answers.
State the match-rate formula, including numerator, denominator, and match definition, for the specific SKU we are evaluating.
Provide the as-of date of the file used for the quoted rate and the recency rule for stale keys.
Separate person, device, and household links. Do not blend them.
State whether modeled or probabilistic links are included in the quoted rate, and how to exclude them.
List each drop between input rows and the denominator you used, with counts.
Name the destination, if any, whose accepted-key count is included in the rate.
Describe seed prep (hash, case, trim, encoding) and confirm you will echo those rules on our file.
Agree to re-run the same seed during the evaluation and explain material changes.
When two finalists answer those questions, you can compare them. Until then, you are ranking decks. For the operational cost of mixing deterministic and probabilistic links in one number, see what deterministic vs probabilistic matching costs. For how to run the file movement and holdout, stay with the seed match testing and pilot pages. Questions on a scoped sample: contact.
Frequently Asked Questions
Can I compare two vendors' published match rates?
Only after both vendors write the same match definition, denominator, recency window, and household rule, and both run the same seed under that spec. Published rates that omit those items are not comparable.
What should count as a match in an activation workflow?
A match that produces a key the named destination will accept, at the relationship level your use requires. Internal graph joins that fail onboarding are not usable matches for that workflow.
How should household expansion appear in the report?
As its own line. Report person or account matches separately from device and household endpoints, and include a capped vs uncapped view if expansion is in scope.
Does a large identity graph imply high match rate on my file?
No. Graph size is not a substitute for a seed test on the cohort, geography, and identifier types you actually activate.
What is the minimum matched-sample report?
Seed echo and prep rules, valid-key count, eligible resolution by relationship type, destination-accepted count, suppression, unmatched with reason codes, and a re-run during the evaluation.
✓ Opt-Out Request Honored via Global Privacy Control
We use cookies and similar technologies to improve your experience. No marketing cookies are pre-selected. Learn more in our Privacy Policy. Your Privacy Choices.