Synthetic Data in Audience Segmentation
Tags: Trends, Statistics
TL;DR
- AI‑driven segmentation is surging: 35% of publishers and 51% of agencies already use AI for this.
- 25–33% of agencies/brands use generative AI to build segments with synthetic data where signals are scarce.
- AI helps rebuild addressability without cross‑site user‑level data, but inference of sensitive traits raises scrutiny.
Why it matters for HK marketers: Use AI to offset signal loss, but police data sources, labels, and synthetic‑data limits.
What’s changing
- Precision at scale: Real‑time and historical attributes power complex segments.
- Signal‑loss workaround: AI generates targeting criteria without regulated health or cross‑site identifiers.
- Synthetic data: Enables modeling/testing when first‑party data is thin or risky (e.g., minors, health), but watch for homogeneity, hallucinations, and validation overhead.
Governance to build in
- Data source diligence: Scraping, RAG, CDP feeds, and third‑party inputs must be contractually clear on consent, deletion, and audit.
- Label reviews: Avoid inferences that reveal emotional state, race, gender, or other sensitive traits.
35% of publishers use AI for segmentation; 51% of agencies do.
One‑quarter to one‑third of agencies and brands use generative AI to build segments with synthetic data.
So what for marketers
Adopt AI segmentation with a label review process and vendor audits. Pilot synthetic data where privacy risk is high, and track ROI versus validation costs.
← Back to Knowledge Base