The keys that connect a device, an email address or a household across datasets, and how to judge a join.
MAID (Mobile Advertising ID)
A MAID is a resettable identifier that a phone operating system assigns to a device for advertising. On iOS it is called the IDFA and on Android the GAID (also written AAID).
Because the user can reset, delete or withdraw permission for it, and the platform returns zeros once permission is gone, a MAID identifies a device for a stretch of time rather than a person for life. MAIDs show up in app SDK data, programmatic bid logs and location panels, which is why they are the most common key for joining mobile datasets to each other.
Example: A retailer takes the MAIDs seen inside its stores last quarter and sends them to a demand-side platform to show those phones a loyalty offer.
How buyers use it: Buyers use MAIDs to onboard mobile audiences, measure store visits after ad exposure and join CTV or location outcomes when the contract allows it. Ask how recently each ID was seen and how resets and opt-outs are removed.
Related: MAID Feed, Identity graph
IDFA (Identifier for Advertisers)
The IDFA is Apple's advertising identifier for iPhone and iPad. Since iOS 14.5 an app only receives a real IDFA after the user grants permission through App Tracking Transparency.
If the user declines, or the app never asks, Apple returns a value of all zeros instead of a usable ID. That makes the IDFA an opt-in identifier, and it is the main reason iOS coverage in any MAID dataset is thinner than Android coverage. The value is a standard UUID written in uppercase hex with hyphens.
Example: A fitness app shows the tracking prompt, the user taps Allow, and the app can now pass that device IDFA to its ad partners for frequency capping.
How buyers use it: Buyers ask vendors to report iOS and Android counts separately and to drop zeroed IDFAs before quoting a match rate, since a string of zeros matches everything and means nothing.
Sources: Apple developer documentation for advertisingIdentifier
Related: MAID Feed
GAID / AAID (Google Advertising ID)
The GAID, also called the AAID or Android Advertising ID, is the resettable advertising identifier that Google Play services provides on Android devices.
Unlike the IDFA it is available by default, but users can reset it or delete it in Settings under Privacy and Ads. Google says that once a user deletes it, apps asking for the identifier receive a string of zeros. Apps targeting recent Android versions also have to declare the AD_ID permission in their manifest and complete the Advertising ID declaration in Play Console.
Example: A mobile game reads the GAID at install, and an attribution partner uses it to credit the install to the video ad the same phone saw the day before.
How buyers use it: Buyers treat GAIDs as the larger share of most US MAID files and check that deleted and zeroed IDs are filtered out and that resets are handled before a file is refreshed.
Sources: Google Play Console Help, Advertising ID
Related: MAID Feed
IDFA vs GAID
IDFA and GAID are both mobile advertising IDs, but the IDFA is opt-in on iOS through App Tracking Transparency while the GAID is available on Android unless the user resets or deletes it.
That one difference drives most of what buyers see in practice. Android usually supplies more usable IDs, iOS IDs are fewer but come with an explicit permission, and the two platforms reset and zero out IDs under different rules. In a data file both are usually stored in one MAID column with a separate field for the operating system.
Example: A MAID file of a million devices might be mostly Android. A campaign aimed at iPhone owners in that file reaches far fewer devices than the headline count suggests.
How buyers use it: Buyers ask for counts, recency and match rates split by platform, and write that split into the pilot acceptance criteria so an Android-heavy file is not sold as balanced reach.
Related: MAID Feed, Identity graph
HEM (Hashed Email)
A HEM is an email address run through a one-way hash function, most often SHA-256, so two parties can match on the same address without exchanging the address itself.
The same normalized email always produces the same hash, which is what makes it a join key. The hash cannot be reversed by math, but it is still personal data under most US state privacy laws because it can be matched back to a person by anyone who holds the plain address. Hashes only match when both sides normalized the email the same way before hashing.
Example: A brand hashes its CRM emails with SHA-256 and sends them to an onboarding partner, which matches them to hashed emails linked to devices and CTV households.
How buyers use it: Buyers use HEMs as the bridge between a CRM file and an identity graph, and they specify the hash algorithm and normalization rules in writing before any match test.
Related: Core Email File, Identity graph
SHA-256 normalization
SHA-256 normalization means cleaning an email or phone number into one agreed format before hashing it with SHA-256, so the same input on both sides produces the same hash.
Typical email rules are to trim spaces and convert to lowercase. Some frameworks go further and, for Gmail addresses, remove dots and anything after a plus sign. Phone numbers are usually put into E.164 format with the country code. Hashing is exact, so one stray capital letter or trailing space gives a completely different 64 character hex string and a lost match.
Example: One side hashes "Jane.Doe@Example.com " and the other hashes "jane.doe@example.com". The hashes share nothing, so a real customer shows up as a non-match.
How buyers use it: Buyers put the normalization steps into the match test instructions and ask the vendor to confirm them, because many low match rates turn out to be formatting problems, not coverage problems.
Related: Core Email File, Seed match testing
UID2 (Unified ID 2.0)
UID2 is an open-source advertising identifier framework, started by The Trade Desk, that turns a consenting user email address or phone number into a pseudonymous ID and an encrypted token for programmatic advertising.
The email or phone number is normalized and hashed, then converted by a UID2 operator into a raw UID2. Publishers pass an encrypted, regularly refreshed token in the bid stream rather than the raw value, and participants agree to the UID2 policies, including honoring opt-outs through the UID2 system.
Example: A news site asks readers to log in, generates UID2 tokens from their emails and passes them to buyers, who match them to UID2s generated from their own CRM.
How buyers use it: Buyers use UID2 to reach known customers on the open web without third-party cookies, and they check whether a data partner can supply or receive UID2s alongside HEMs and MAIDs.
Sources: The Trade Desk, Unified ID 2.0
Related: Identity resolution
RampID
RampID is LiveRamp's pseudonymous, person-based identifier. It links a person's known identifiers, such as emails and postal addresses, into one ID without passing the underlying personal data to the brands and platforms that use it.
LiveRamp encodes RampIDs differently for each partner, so two companies cannot simply compare their raw RampIDs. Matching happens inside the LiveRamp network. That design protects the underlying data and also means RampID works best when both sides of a join already use LiveRamp.
Example: A retailer onboards its loyalty file through LiveRamp, receives RampIDs, and activates them on a streaming TV platform that also connects through LiveRamp.
How buyers use it: Buyers ask whether a vendor can deliver or accept RampIDs, or whether the join will run on hashed emails or MAIDs instead, since that decides which platforms the audience can reach.
Sources: LiveRamp developer docs, About RampIDs
Related: GSDSI vs LiveRamp, Identity resolution
ID5 ID
The ID5 ID is a shared identifier from the company ID5 that publishers and ad platforms use to recognize users across sites without third-party cookies.
The publisher shares signals such as a hashed email when a user logs in, and ID5 returns an encrypted ID to the page. Only platforms authorized under the user and publisher privacy choices can decrypt it to a stable value. Because it can work without a login, it tends to cover more traffic than email-only IDs, with less certainty about who each ID is.
Example: A European publisher adds the ID5 module to its header bidding setup, and buyers bidding on its pages receive an ID5 ID they can use for frequency capping.
How buyers use it: Buyers list which shared IDs (ID5, UID2, RampID) each platform in a plan supports, then check which of those a data partner can map to, before promising addressable reach.
Sources: ID5 documentation for publishers
Related: Identity resolution
Device graph
A device graph is a dataset that links identifiers believed to belong to the same person or household, such as MAIDs, hashed emails, CTV IDs and IP addresses.
Each link is an edge, and each edge comes from evidence such as a login seen on two devices or two devices sharing a home network overnight. Some edges are deterministic and some are probabilistic, and good graphs label which is which, with a first seen and last seen date on every edge.
Example: A graph links a hashed email, two phones and a smart TV to one household, so an ad seen on the TV can be tied to a store visit recorded on one of the phones.
How buyers use it: Buyers use a device graph to extend a CRM audience to more screens and to connect exposure to outcomes. They ask for edge counts by type, recency fields and how the graph handles shared devices.
Related: Identity graph, MAID Feed
Device graph decay
Device graph decay is the rate at which links in an identity graph go stale as people replace phones, reset advertising IDs, move house or change email addresses.
An edge that was true six months ago may now point at a device someone sold. Decay is why last seen dates matter more than total edge counts. Vendors manage it with refresh cycles, change files that add and remove edges, and rules that expire links nobody has seen for a set period.
Example: A household graph built last year still links a family to an old tablet. Ads keep going to that tablet, and none of the resulting visits can be credited.
How buyers use it: Buyers filter on last seen fields during warehouse merges, ask how often edges are removed as well as added, and compare match rates from a fresh pull against an older copy.
Related: Device graph decay resource, Identity graph
Identity resolution
Identity resolution is the process of deciding which records, devices and identifiers in different datasets refer to the same person, household or business.
It combines matching rules, reference data such as an identity graph, and decisions about what to do when evidence conflicts. The output is usually a persistent ID attached to each input record, plus a confidence level or match method, so later systems can tell a strong link from a guess.
Example: A bank has one customer listed three times with slightly different names and addresses. Identity resolution collapses those rows into one person and attaches her phone and hashed email.
How buyers use it: Buyers use identity resolution to clean CRM files, build audiences that reach more channels and measure outcomes across devices. They ask vendors to report results by match method, not only as one total.
Related: Identity resolution buying guide, Identity graph
Deterministic vs probabilistic matching
Deterministic matching links records that share an exact identifier, such as the same hashed email. Probabilistic matching links records that are statistically likely to belong together, based on signals like IP address, device type and timing.
Deterministic links are more precise but cover fewer records. Probabilistic links reach further but some will be wrong, and how many depends on the model and threshold. Neither is better in general. The mistake is blending them into one number without saying which is which.
Example: A phone and a laptop that logged into the same account are a deterministic match. Two devices that share a home network every night but never log in are a probabilistic match.
How buyers use it: Buyers ask for match results split by method and set different uses for each, for example deterministic links only for measurement and both types for broad reach campaigns.
Related: Deterministic vs probabilistic cost, Identity resolution
Match rate
Match rate is the share of records in your file that a vendor can link to its own data on the agreed keys. It depends as much on your file and the match rules as on the vendor's coverage.
The number changes with the seed file, the keys used, the date of the vendor data, and whether probabilistic links count. It should be stated with a numerator, a denominator and a date. A vendor headline figure measured on someone else’s file tells you little about yours.
Example: A brand sends a seed of hashed emails. The vendor matches a portion of them to at least one MAID seen in the last ninety days. That portion, on that date, is the match rate for that use.
How buyers use it: Buyers measure match rate on their own seed, split by geography and cohort, and write the definition into the pilot charter before any file is exchanged.
Related: Pilot process, Why match rates are not comparable