Skip to content

Data glossary

Plain definitions for 44 terms buyers run into when licensing identity, location, TV and audience data. Each entry starts with a short definition, then gives an example and how buyers use it. Link to any term with its anchor, for example /glossary#maid.

Identity and matching

The keys that connect a device, an email address or a household across datasets, and how to judge a join.

MAID (Mobile Advertising ID)IDFA (Identifier for Advertisers)GAID / AAID (Google Advertising ID)IDFA vs GAIDHEM (Hashed Email)SHA-256 normalizationUID2 (Unified ID 2.0)RampIDID5 IDDevice graphDevice graph decayIdentity resolutionDeterministic vs probabilistic matchingMatch rate

MAID (Mobile Advertising ID)

A MAID is a resettable identifier that a phone operating system assigns to a device for advertising. On iOS it is called the IDFA and on Android the GAID (also written AAID).

Because the user can reset, delete or withdraw permission for it, and the platform returns zeros once permission is gone, a MAID identifies a device for a stretch of time rather than a person for life. MAIDs show up in app SDK data, programmatic bid logs and location panels, which is why they are the most common key for joining mobile datasets to each other.

Example: A retailer takes the MAIDs seen inside its stores last quarter and sends them to a demand-side platform to show those phones a loyalty offer.

How buyers use it: Buyers use MAIDs to onboard mobile audiences, measure store visits after ad exposure and join CTV or location outcomes when the contract allows it. Ask how recently each ID was seen and how resets and opt-outs are removed.

Related: MAID Feed, Identity graph

IDFA (Identifier for Advertisers)

The IDFA is Apple's advertising identifier for iPhone and iPad. Since iOS 14.5 an app only receives a real IDFA after the user grants permission through App Tracking Transparency.

If the user declines, or the app never asks, Apple returns a value of all zeros instead of a usable ID. That makes the IDFA an opt-in identifier, and it is the main reason iOS coverage in any MAID dataset is thinner than Android coverage. The value is a standard UUID written in uppercase hex with hyphens.

Example: A fitness app shows the tracking prompt, the user taps Allow, and the app can now pass that device IDFA to its ad partners for frequency capping.

How buyers use it: Buyers ask vendors to report iOS and Android counts separately and to drop zeroed IDFAs before quoting a match rate, since a string of zeros matches everything and means nothing.

Sources: Apple developer documentation for advertisingIdentifier

Related: MAID Feed

GAID / AAID (Google Advertising ID)

The GAID, also called the AAID or Android Advertising ID, is the resettable advertising identifier that Google Play services provides on Android devices.

Unlike the IDFA it is available by default, but users can reset it or delete it in Settings under Privacy and Ads. Google says that once a user deletes it, apps asking for the identifier receive a string of zeros. Apps targeting recent Android versions also have to declare the AD_ID permission in their manifest and complete the Advertising ID declaration in Play Console.

Example: A mobile game reads the GAID at install, and an attribution partner uses it to credit the install to the video ad the same phone saw the day before.

How buyers use it: Buyers treat GAIDs as the larger share of most US MAID files and check that deleted and zeroed IDs are filtered out and that resets are handled before a file is refreshed.

Sources: Google Play Console Help, Advertising ID

Related: MAID Feed

IDFA vs GAID

IDFA and GAID are both mobile advertising IDs, but the IDFA is opt-in on iOS through App Tracking Transparency while the GAID is available on Android unless the user resets or deletes it.

That one difference drives most of what buyers see in practice. Android usually supplies more usable IDs, iOS IDs are fewer but come with an explicit permission, and the two platforms reset and zero out IDs under different rules. In a data file both are usually stored in one MAID column with a separate field for the operating system.

Example: A MAID file of a million devices might be mostly Android. A campaign aimed at iPhone owners in that file reaches far fewer devices than the headline count suggests.

How buyers use it: Buyers ask for counts, recency and match rates split by platform, and write that split into the pilot acceptance criteria so an Android-heavy file is not sold as balanced reach.

Related: MAID Feed, Identity graph

HEM (Hashed Email)

A HEM is an email address run through a one-way hash function, most often SHA-256, so two parties can match on the same address without exchanging the address itself.

The same normalized email always produces the same hash, which is what makes it a join key. The hash cannot be reversed by math, but it is still personal data under most US state privacy laws because it can be matched back to a person by anyone who holds the plain address. Hashes only match when both sides normalized the email the same way before hashing.

Example: A brand hashes its CRM emails with SHA-256 and sends them to an onboarding partner, which matches them to hashed emails linked to devices and CTV households.

How buyers use it: Buyers use HEMs as the bridge between a CRM file and an identity graph, and they specify the hash algorithm and normalization rules in writing before any match test.

Related: Core Email File, Identity graph

SHA-256 normalization

SHA-256 normalization means cleaning an email or phone number into one agreed format before hashing it with SHA-256, so the same input on both sides produces the same hash.

Typical email rules are to trim spaces and convert to lowercase. Some frameworks go further and, for Gmail addresses, remove dots and anything after a plus sign. Phone numbers are usually put into E.164 format with the country code. Hashing is exact, so one stray capital letter or trailing space gives a completely different 64 character hex string and a lost match.

Example: One side hashes "Jane.Doe@Example.com " and the other hashes "jane.doe@example.com". The hashes share nothing, so a real customer shows up as a non-match.

How buyers use it: Buyers put the normalization steps into the match test instructions and ask the vendor to confirm them, because many low match rates turn out to be formatting problems, not coverage problems.

Related: Core Email File, Seed match testing

UID2 (Unified ID 2.0)

UID2 is an open-source advertising identifier framework, started by The Trade Desk, that turns a consenting user email address or phone number into a pseudonymous ID and an encrypted token for programmatic advertising.

The email or phone number is normalized and hashed, then converted by a UID2 operator into a raw UID2. Publishers pass an encrypted, regularly refreshed token in the bid stream rather than the raw value, and participants agree to the UID2 policies, including honoring opt-outs through the UID2 system.

Example: A news site asks readers to log in, generates UID2 tokens from their emails and passes them to buyers, who match them to UID2s generated from their own CRM.

How buyers use it: Buyers use UID2 to reach known customers on the open web without third-party cookies, and they check whether a data partner can supply or receive UID2s alongside HEMs and MAIDs.

Sources: The Trade Desk, Unified ID 2.0

Related: Identity resolution

RampID

RampID is LiveRamp's pseudonymous, person-based identifier. It links a person's known identifiers, such as emails and postal addresses, into one ID without passing the underlying personal data to the brands and platforms that use it.

LiveRamp encodes RampIDs differently for each partner, so two companies cannot simply compare their raw RampIDs. Matching happens inside the LiveRamp network. That design protects the underlying data and also means RampID works best when both sides of a join already use LiveRamp.

Example: A retailer onboards its loyalty file through LiveRamp, receives RampIDs, and activates them on a streaming TV platform that also connects through LiveRamp.

How buyers use it: Buyers ask whether a vendor can deliver or accept RampIDs, or whether the join will run on hashed emails or MAIDs instead, since that decides which platforms the audience can reach.

Sources: LiveRamp developer docs, About RampIDs

Related: GSDSI vs LiveRamp, Identity resolution

ID5 ID

The ID5 ID is a shared identifier from the company ID5 that publishers and ad platforms use to recognize users across sites without third-party cookies.

The publisher shares signals such as a hashed email when a user logs in, and ID5 returns an encrypted ID to the page. Only platforms authorized under the user and publisher privacy choices can decrypt it to a stable value. Because it can work without a login, it tends to cover more traffic than email-only IDs, with less certainty about who each ID is.

Example: A European publisher adds the ID5 module to its header bidding setup, and buyers bidding on its pages receive an ID5 ID they can use for frequency capping.

How buyers use it: Buyers list which shared IDs (ID5, UID2, RampID) each platform in a plan supports, then check which of those a data partner can map to, before promising addressable reach.

Sources: ID5 documentation for publishers

Related: Identity resolution

Device graph

A device graph is a dataset that links identifiers believed to belong to the same person or household, such as MAIDs, hashed emails, CTV IDs and IP addresses.

Each link is an edge, and each edge comes from evidence such as a login seen on two devices or two devices sharing a home network overnight. Some edges are deterministic and some are probabilistic, and good graphs label which is which, with a first seen and last seen date on every edge.

Example: A graph links a hashed email, two phones and a smart TV to one household, so an ad seen on the TV can be tied to a store visit recorded on one of the phones.

How buyers use it: Buyers use a device graph to extend a CRM audience to more screens and to connect exposure to outcomes. They ask for edge counts by type, recency fields and how the graph handles shared devices.

Related: Identity graph, MAID Feed

Device graph decay

Device graph decay is the rate at which links in an identity graph go stale as people replace phones, reset advertising IDs, move house or change email addresses.

An edge that was true six months ago may now point at a device someone sold. Decay is why last seen dates matter more than total edge counts. Vendors manage it with refresh cycles, change files that add and remove edges, and rules that expire links nobody has seen for a set period.

Example: A household graph built last year still links a family to an old tablet. Ads keep going to that tablet, and none of the resulting visits can be credited.

How buyers use it: Buyers filter on last seen fields during warehouse merges, ask how often edges are removed as well as added, and compare match rates from a fresh pull against an older copy.

Related: Device graph decay resource, Identity graph

Identity resolution

Identity resolution is the process of deciding which records, devices and identifiers in different datasets refer to the same person, household or business.

It combines matching rules, reference data such as an identity graph, and decisions about what to do when evidence conflicts. The output is usually a persistent ID attached to each input record, plus a confidence level or match method, so later systems can tell a strong link from a guess.

Example: A bank has one customer listed three times with slightly different names and addresses. Identity resolution collapses those rows into one person and attaches her phone and hashed email.

How buyers use it: Buyers use identity resolution to clean CRM files, build audiences that reach more channels and measure outcomes across devices. They ask vendors to report results by match method, not only as one total.

Related: Identity resolution buying guide, Identity graph

Deterministic vs probabilistic matching

Deterministic matching links records that share an exact identifier, such as the same hashed email. Probabilistic matching links records that are statistically likely to belong together, based on signals like IP address, device type and timing.

Deterministic links are more precise but cover fewer records. Probabilistic links reach further but some will be wrong, and how many depends on the model and threshold. Neither is better in general. The mistake is blending them into one number without saying which is which.

Example: A phone and a laptop that logged into the same account are a deterministic match. Two devices that share a home network every night but never log in are a probabilistic match.

How buyers use it: Buyers ask for match results split by method and set different uses for each, for example deterministic links only for measurement and both types for broad reach campaigns.

Related: Deterministic vs probabilistic cost, Identity resolution

Match rate

Match rate is the share of records in your file that a vendor can link to its own data on the agreed keys. It depends as much on your file and the match rules as on the vendor's coverage.

The number changes with the seed file, the keys used, the date of the vendor data, and whether probabilistic links count. It should be stated with a numerator, a denominator and a date. A vendor headline figure measured on someone else’s file tells you little about yours.

Example: A brand sends a seed of hashed emails. The vendor matches a portion of them to at least one MAID seen in the last ninety days. That portion, on that date, is the match rate for that use.

How buyers use it: Buyers measure match rate on their own seed, split by geography and cohort, and write the definition into the pilot charter before any file is exchanged.

Related: Pilot process, Why match rates are not comparable

Location and places

How places are drawn, how visits are counted, and where the law draws lines around location data.

POI (Point of Interest)POI polygonCentroidGeofenceDwell timeFoot trafficVisitation panelPrecise geolocationSensitive location

POI (Point of Interest)

A POI is a real place that people visit, such as a store, restaurant, stadium or office, stored with its location, category, brand and address.

A POI record can hold a single point, a building outline, or both. The metadata matters as much as the geometry, since brand, category, opening status and parent chain are what turn raw locations into business questions. POI datasets also need regular updates because stores open, close and move constantly.

Example: A coffee chain location is stored with its brand, category, street address, a point at the front door and a polygon around the building.

How buyers use it: Buyers use POI data to build geofences, count visits, plan new sites and label location data with brands and categories. They check coverage for their own brands first.

Related: POI data, POI geofencing

POI polygon

A POI polygon is the drawn outline of a place, usually its building footprint or property boundary, rather than a single point on a map.

Polygons let you count a visit only when a device is actually inside the place, which matters in malls, strip centers and dense city blocks where a circle around a point would catch the store next door. Quality depends on how the shapes were drawn and how often they are checked against the real world.

Example: A pharmacy and a burger restaurant share a wall. Radius geofences around each would overlap, while polygons keep their visits apart.

How buyers use it: Buyers ask for polygon coverage by brand and category, and use polygons for visit counting and attribution while keeping points for maps and rough reach.

Related: POI geofencing, Polygons vs radius

Centroid

A centroid is a single latitude and longitude point that represents a place, usually the geometric center of its shape or a point placed at the building or front door.

Centroids are compact and easy to work with, which is why most address and POI files include them. The weakness is that a point says nothing about how big a place is. A centroid for a large retailer and one for a small kiosk look the same, and a radius drawn around either may include other businesses.

Example: A big box store centroid sits in the middle of its roof. A radius around it misses shoppers in the far corner of the parking lot and catches the gas station next door.

How buyers use it: Buyers accept centroids for mapping, routing and rough targeting, but ask for polygons when visits or foot traffic will be counted and reported.

Related: POI data, Point of interest data guide

Geofence

A geofence is a virtual boundary around a real place, drawn as a radius or a polygon, that triggers an action or a count when a device enters or leaves it.

In advertising, geofences decide who joins an audience, such as everyone seen at car dealerships last month, and when a visit counts for measurement. The size and shape of the fence decide accuracy. Some state health privacy laws restrict geofences drawn around health care facilities, so the place list matters as much as the geometry.

Example: An auto brand geofences competitor dealerships and builds an audience of devices seen inside them during the last thirty days.

How buyers use it: Buyers choose polygon or radius fences depending on the question, exclude sensitive places, and ask vendors which place list and version the fences were built from.

Related: POI geofencing, Polygons vs radius

Dwell time

Dwell time is how long a device stays inside a place during one visit, measured from the first location signal inside the boundary to the last.

It separates real visits from devices passing by and helps tell a quick pickup from a long shopping trip. Because phones report location at uneven intervals, dwell time is an estimate, and vendors use minimum and maximum dwell rules to decide what counts as a visit at all.

Example: A device seen inside a grocery store for twenty-five minutes counts as a shopping visit. A device seen for thirty seconds while driving past on the road does not.

How buyers use it: Buyers ask for the dwell rules behind each visit count and keep them fixed for the length of a test, since changing the minimum dwell changes the visit totals.

Related: Global mobility data, Location intelligence

Foot traffic

Foot traffic is the estimated number of visits to a place or group of places over a period of time, built from location data and scaled to represent the wider population.

Raw device counts only show the phones in a panel, so providers adjust them for panel size, demographics and geography to estimate true visits. The trend over time is usually more reliable than any single total, and results depend on the place boundaries and visit rules used.

Example: An analyst compares weekly visits to two restaurant chains in the same markets to see which one gained share after a menu launch.

How buyers use it: Buyers use foot traffic for site selection, competitive benchmarking and investment research, and ask how the panel is scaled and how often the method changes.

Related: Competitive benchmarking, Location intelligence

Visitation panel

A visitation panel is the set of devices whose location data a provider uses to measure visits, tracked consistently enough over time to support trend analysis.

The panel is a sample, not a census. Its size, stability and mix of regions and people decide how far results can be trusted. A panel that loses a large data source overnight can create a fake drop in visits, which is why good providers publish panel counts and adjust for changes.

Example: A hedge fund tracks weekly visits to a retailer. The panel shrinks after an app partner leaves, so the provider rebalances the panel and restates the history.

How buyers use it: Buyers ask for panel size over time, how the panel is balanced and how breaks in the history are flagged, especially when the data feeds trading or forecasting models.

Related: Global mobility data, Alternative data for finance

Precise geolocation

Precise geolocation is location data accurate enough to place a person within a small circle. Several US state privacy laws set that circle at a radius of 1,750 feet, and California sets it at 1,850 feet.

Virginia's Consumer Data Protection Act defines precise geolocation data as information that identifies a person's location within a radius of 1,750 feet, and several other state laws use the same figure. The California Consumer Privacy Act uses a circle with a radius of 1,850 feet. Under these laws precise geolocation is usually sensitive data, which can mean opt-in consent or a right to limit its use, and in 2026 Virginia amended its law to bar the sale of it, joining Maryland and Oregon.

Example: A ZIP code is coarse location. A GPS reading that puts a phone inside a specific store is precise geolocation under these definitions.

How buyers use it: Buyers ask whether a feed contains precise geolocation, which states it covers, what consent was collected, and whether they need coarser location for the use they have in mind.

Sources: Virginia Code 59.1-575 (definitions), California Civil Code 1798.140 (definitions), Virginia SB 338 sale ban (Troutman Pepper Locke)

Related: Privacy and compliance, Global mobility data

Sensitive location

A sensitive location is a place whose visits can reveal private facts about a person, such as medical facilities, places of worship, shelters, schools, correctional facilities and labor union offices.

The FTC orders against X-Mode (now Outlogic) and other location data companies barred them from selling data about visits to places like these and required them to keep a list of such places. Several states also restrict geofences around health care facilities. There is no single national list, so each provider keeps its own and should document it.

Example: A mobility feed removes every device ping inside a hospital polygon before delivery, and the manifest records which version of the sensitive place list was applied.

How buyers use it: Buyers ask for the exclusion list, how it is updated and how it is applied, and add a contract clause barring any attempt to rebuild visits to excluded places.

Sources: FTC, final order against X-Mode and Outlogic

Related: Privacy and compliance, Sensitive location checklist

TV, audience and data types

Viewing data, where audience data comes from, and the environments used to join it safely.

ACR (Automatic Content Recognition)CTV (Connected TV)First-party dataSecond-party dataThird-party dataData clean roomData onboardingLookalike modelingSuppression fileOpt-out and consent signalsAlternative data

ACR (Automatic Content Recognition)

ACR is technology built into many smart TVs that identifies what is on screen by matching small samples of the picture or sound against a reference library.

It works whatever the source, including cable, streaming apps, game consoles and antennas, which is why it shows viewing that ad server logs miss. TV makers usually show a viewing data prompt during setup, and state enforcement since 2025 has pushed them toward clearer opt-in consent. In Texas, Samsung and LG settled with the attorney general in 2026 and agreed to collect ACR data only with express consent, so ask how consent was collected on the TVs behind a feed. The output is a log of what played on which TV and when, tied to a device or household ID rather than a named person.

Example: A TV logs that a household watched a national ad during a football game, even though the game came through a cable box and not an app.

How buyers use it: Buyers use ACR data to measure CTV and linear reach and frequency, find households that saw an ad, and connect those exposures to site visits or sales.

Sources: IAPP, ACR technology takes the privacy enforcement spotlight

Related: CTV and ACR data, CTV and ACR hub

CTV (Connected TV)

CTV means a television connected to the internet, either a smart TV or a TV using a streaming device or game console, and the ads delivered through its streaming apps.

CTV ads are bought programmatically or directly from streaming services and can be targeted to households with data. Because the TV is a shared screen, CTV works at the household level, and it is often measured with IP addresses, device graphs and ACR data.

Example: A streaming service serves a car ad to households in a set of ZIP codes, and the advertiser later checks how many of those households visited a dealership.

How buyers use it: Buyers use CTV to reach cord cutters, extend TV campaigns and run outcome studies. They ask how CTV IDs are linked to households and how often those links are refreshed.

Related: CTV and ACR hub, Cross-channel measurement

First-party data

First-party data is information a company collects directly from its own customers and users, such as purchases, account details, website activity and loyalty records.

Because it comes straight from the customer relationship, a company can usually trust it more than data about the same people bought from outside, and its use is governed by the company privacy notice and the consents it collected. Its limit is reach, since it only covers people who already deal with you, which is why companies add second- and third-party data.

Example: An airline uses its loyalty members' trip history to decide which of them get a credit card offer.

How buyers use it: Buyers use first-party data as the seed for matching, lookalike modeling and measurement, and check that their own privacy notice covers sharing it with an onboarding partner.

Related: Audience targeting, Identity resolution

Second-party data

Second-party data is another company's first-party data that it shares directly with a known partner under an agreement, rather than selling it on the open market.

Because the source is a known partner, the buyer can see how it was collected. The sharing usually happens in a clean room or through a direct file transfer with contract limits on use. The value depends on how well the partner customers overlap with yours.

Example: A grocery chain shares purchase data with a food brand inside a clean room so the brand can see which of its ads led to sales in those stores.

How buyers use it: Buyers use second-party data for joint measurement and co-marketing, and they write permitted use, retention and aggregation rules into the data sharing agreement.

Related: Clean room measurement, Cross-channel measurement

Third-party data

Third-party data is information compiled by a company that has no direct relationship with the people in it, then licensed to other businesses.

It is how most companies reach beyond their own customers, filling gaps in demographics, interests, location and identity. Businesses that sell it are data brokers under several state laws. Quality varies a great deal, so provenance, consent chain and refresh cadence matter more than headline counts. Two vendors can describe the same segment in the same words and still deliver very different files, which is why a matched sample beats a spec sheet.

Example: A home insurer licenses household data to find homeowners in its service area who are not yet customers.

How buyers use it: Buyers use third-party data for prospecting, enrichment and modeling, and ask each vendor where the data came from, what notice was given and which uses are permitted.

Related: Products, Sourcing methodology

Data clean room

A data clean room is a controlled environment where two or more parties can join their data and run approved analyses without either side seeing the other’s raw records.

Clean rooms enforce rules on what can be queried and what comes out, such as minimum group sizes and aggregate-only results. They are offered by cloud warehouses, ad platforms and specialist vendors. A clean room controls how the join runs, but it does not replace consent, exclusions or contract terms for the data inside it.

Example: A streaming service and an automaker join ad exposure and dealer visit data in a clean room and export only household counts by campaign.

How buyers use it: Buyers use clean rooms for match tests before licensing and for measurement after a campaign. They confirm the minimum group size and who can approve new queries.

Related: Clean room measurement, Clean rooms in 2026

Data onboarding

Data onboarding is the process of taking offline or CRM records, such as names, emails and addresses, and matching them to online identifiers so they can be used in digital advertising.

The onboarder normalizes and hashes the input, matches it to an identity graph and passes the matched IDs to ad platforms, usually without sending the personal data itself. Every step loses some records, so the size of the audience that arrives in a platform is smaller than the file that went in.

Example: A car dealer uploads last year’s service customers, and the onboarder matches them to MAIDs and CTV IDs for a reminder campaign.

How buyers use it: Buyers track how many records survive each step of onboarding and ask which identifiers are sent to which platforms, so they can see where their audience shrinks.

Related: Audience targeting, Integrations

Lookalike modeling

Lookalike modeling finds new people or households who resemble a seed audience, such as your best customers, by scoring a larger population on the traits that set the seed apart.

The model learns which attributes are more common in the seed than in the general population, then ranks everyone else by how closely they match. A bigger audience is usually less similar to the seed, so buyers choose a point on that tradeoff. A biased or tiny seed produces a biased lookalike.

Example: A subscription box company models its most loyal subscribers and targets the top ranked households that are not yet customers.

How buyers use it: Buyers use lookalikes for prospecting, test several audience sizes against a holdout group, and check that the model does not use traits they are not allowed to target on.

Related: Audience targeting, Consumer data file

Suppression file

A suppression file is a list of people, emails, devices or addresses that must be removed from a dataset or campaign before it is used.

Suppression lists come from consumer opt-outs, deletion requests, do-not-contact registries, existing customers and legal exclusions. They only work if they are applied at every step and kept current, so a suppression file applied once and never refreshed quickly falls out of date. The buyer usually has its own list too, and the contract should say whose list wins when the two disagree.

Example: Before an email campaign, a brand removes everyone on its unsubscribe list and everyone who asked a data vendor to delete their information.

How buyers use it: Buyers ask how a vendor applies opt-outs and deletions to feeds already delivered, how fast updates arrive, and how to apply their own suppression list before activation.

Related: Privacy and compliance, Core Email File

Opt-out and consent signals

Opt-out and consent signals are the records of what a person has agreed to or refused, such as an app tracking permission, a cookie banner choice or a browser opt-out preference signal like Global Privacy Control.

California privacy regulations require businesses to treat a valid opt-out preference signal, such as Global Privacy Control, as a request to opt out of the sale or sharing of personal information, and several other states have similar rules. Industry frameworks pass consent strings along the ad supply chain so later parties know what was allowed.

Example: A browser sends the Global Privacy Control header, and the site stops sharing that visitor’s data with advertising partners.

How buyers use it: Buyers ask vendors which signals they capture and how an opt-out travels into files already licensed, then check that their own systems honor the same signals.

Sources: California Privacy Protection Agency, CCPA regulations, Global Privacy Control

Related: Privacy center, Do not sell or share

Alternative data

Alternative data is information from outside traditional financial sources, such as foot traffic, web activity, app usage or consumer purchase panels, used by investors and strategy teams to understand companies and markets.

Much of its value comes from speed, since it can show a trend before a company reports it. For trading use, buyers care about a long, consistent history, mapping to tickers and brands, and controls that keep material nonpublic information and personal data out of the feed.

Example: An investment team tracks weekly visits to a restaurant chain and compares them with the chain’s reported same-store sales.

How buyers use it: Buyers backtest alternative data against reported results, check how the history was built, and ask about delivery delay and ticker mapping before using it in a model.

Related: Alternative data hub, Tickerized data

Licensing, delivery and compliance

The contract, file and regulatory vocabulary that decides what you can actually do with a feed.

Permitted useData licensing agreementRefresh cadenceDelivery formatsSnowflake secure shareSample fileMatched-sample pilotData dictionaryData broker registrationCCPA / CPRA

Permitted use

Permitted use is the list of purposes a data license allows, such as ad targeting, measurement, analytics, fraud prevention or research. Anything outside that list is a breach even when the data is technically in hand.

Permitted use is often narrower than people assume. A file licensed for marketing may not be used for eligibility decisions about credit, insurance, employment or housing, which fall under the Fair Credit Reporting Act. Resale, sharing with affiliates and model training are also often limited or excluded.

Example: A lender licenses consumer data for direct mail marketing. Using the same data to approve or decline loan applications would fall outside permitted use.

How buyers use it: Buyers list every intended use before signing, map each one to a clause in the license, and tell internal teams in writing what the data may not be used for.

Related: Sourcing methodology, Terms

Data licensing agreement

A data licensing agreement is the contract that grants a buyer the right to use a dataset on stated terms, without transferring ownership of the data.

It normally covers permitted use, term and renewal, delivery, refresh commitments, fees, confidentiality, security, what happens to the data when the license ends, compliance duties on both sides and liability. A data processing agreement is often attached when personal data is involved.

Example: A retailer signs a twelve month license for POI and foot traffic data, with monthly refreshes, internal analytics as the only permitted use and deletion of all copies at the end of the term.

How buyers use it: Buyers negotiate permitted use, deletion at term end, refresh remedies and indemnities, and make sure the order form matches the product they tested in the pilot.

Related: License red flags, Data processing agreement

Refresh cadence

Refresh cadence is how often a data feed is updated and delivered, such as daily, weekly, monthly or quarterly.

Cadence should fit how fast the underlying facts change and how fast you need to act on them. Location and bid stream data decay quickly, while property and business records change more slowly. A refresh can be a full file or a change file that only lists additions, updates and removals.

Example: A campaign team gets a weekly MAID refresh so audiences built on Monday do not include devices that went quiet last month.

How buyers use it: Buyers write the refresh cadence and delivery window into the contract as a service level, with a remedy if a refresh is late, and match it to their campaign and reporting schedule.

Related: Refresh cadence and drift, Developers

Delivery formats

Delivery formats describe how a licensed dataset reaches the buyer, meaning both the file format, such as Parquet, CSV or JSON, and the channel, such as Amazon S3, SFTP, a Snowflake share or an API.

Parquet is a compressed column format that loads quickly into warehouses and keeps data types. CSV is universal but loses types and can break on stray commas. APIs suit lookups on demand, while bulk files suit full refreshes. GSDSI deliveries ship as CSV, JSON or Parquet through S3, SFTP, Snowflake Secure Share or API, depending on the feed.

Example: A data team asks for Parquet files in its own S3 bucket, split by date, so each weekly refresh loads straight into its warehouse.

How buyers use it: Buyers settle format, channel, file naming, partitioning and manifests during the pilot so production files load without new engineering work.

Related: Developers, Integrations

Snowflake secure share

A Snowflake secure share is a way to deliver data inside Snowflake, where the provider grants the buyer account live, read-only access to tables without copying files.

The buyer queries the shared tables as if they were local, and updates from the provider appear without a new load. Access can be revoked at the end of the license. Because nothing is copied, storage stays with the provider, but the buyer still needs its own Snowflake account and compute.

Example: A retailer analytics team queries a shared POI table next to its own store sales tables and joins them in one SQL statement.

How buyers use it: Buyers choose secure shares for fast evaluations and warehouse-first teams, and confirm the region, refresh timing and what happens to derived tables when the share ends.

Related: Developers, Integrations

Sample file

A sample file is a small extract of a dataset that a vendor shares so a buyer can check the fields, format and data quality before licensing.

A generic sample shows the schema and typical values. It does not show coverage of your own customers or markets, which needs a matched sample or a seed match test. Samples often carry their own short license limiting use to evaluation.

Example: A buyer receives a thousand-row POI sample for one metro area and checks that brand names, categories and polygons look right before asking for a full match test.

How buyers use it: Buyers use sample files to check the data dictionary against real values, test loading into their systems and spot obvious gaps before investing time in a pilot.

Related: Developers and schemas, Pilot process

Matched-sample pilot

A matched-sample pilot is a scoped test where the vendor matches its data to a seed file from the buyer and returns the joined sample, so match rate, freshness and schema can be checked on the buyer’s own records.

The pilot should have a written charter covering the seed, the keys, the date of the vendor data, the success thresholds and who decides. Matching usually runs on hashed identifiers or inside a clean room, and the results show how the data will perform in production far better than a generic sample does.

Example: A brand sends hashed emails for a set of customers, the vendor returns matched MAIDs with last seen dates, and the brand checks results by region against its acceptance criteria.

How buyers use it: Buyers set pass and fail criteria before the files move and keep the pilot data under the same security rules they will use in production.

Related: Pilot process, Seed match testing

Data dictionary

A data dictionary is the document that lists every field in a dataset with its name, meaning, data type, allowed values and how it is populated.

It is the reference engineering, legal and analytics teams share. A good one flags which fields are personal data, which are modeled rather than observed, and when fields were added or deprecated. Without it, teams guess what a field means and those guesses end up in reports.

Example: The dictionary for a mobility feed explains that horizontal_accuracy is in meters, that the value can be missing, and that rows with poor accuracy are removed before delivery.

How buyers use it: Buyers check the dictionary against the sample file, map each field to a permitted use, and ask for notice before any field changes meaning or is removed.

Related: Developers, QC a data file

Data broker registration

A data broker is a business that collects and sells personal information about consumers it has no direct relationship with, and several US states require such businesses to register every year.

California, Vermont, Texas and Oregon run public registries, and in 2026 Connecticut and New Jersey enacted broker registration laws of their own. California registration now sits under the Delete Act, which also created a single deletion request system for consumers. Registration is a disclosure duty. It does not approve any particular data use.

Example: A company that licenses consumer data it did not collect from its own customers registers with the California Privacy Protection Agency by January 31 and appears on the public registry.

How buyers use it: Buyers check a vendor on each relevant state registry, ask how deletion requests reach licensed files, and keep proof of that check in the vendor file.

Sources: California Delete Act statute (CPPA), Vermont Secretary of State, data brokers, Connecticut Public Act 26-64, Hunton, New Jersey data broker registration law

Related: GSDSI broker registrations, State registration tracker

CCPA / CPRA

The CCPA is California’s consumer privacy law, and the CPRA is the 2020 ballot measure that amended and expanded it. Together they give California residents rights to know, delete, correct and opt out of the sale or sharing of their personal information.

The CPRA created the California Privacy Protection Agency, added a category of sensitive personal information that includes precise geolocation, and extended opt-out rights to sharing for cross-context behavioral advertising. Businesses must honor opt-out preference signals under the regulations.

Example: A California resident asks a data company to delete her data, and the company removes her from its files and from the next refresh sent to its clients.

How buyers use it: Buyers write into the license how opt-outs and deletions flow into delivered files, how fast, and who handles consumer requests that arrive at the buyer instead of the vendor.

Sources: California Privacy Protection Agency

Related: Privacy policy, Do not sell or share

For how procurement teams put these terms into RFPs and contracts, see the procurement checklist for data RFPs. For the buying process itself see the FAQ, pilot process, and vendor comparisons.

This glossary is buyer education, not legal advice. Laws change and differ by state, so confirm current requirements with counsel for your use case.

✓ Opt-Out Request Honored via Global Privacy Control