Synthetic personas now power social simulations, market research, policy analysis, and model evaluation. But a profile that sounds vivid is not necessarily realistic, representative, or faithfully enacted by a language model.
Our position paper, accepted to the COLM 2026 Workshop on Social Simulation with LLMs, argues that synthetic personas need explicit grounding. Researchers should state where persona information comes from, who or what the personas are intended to represent, what evidence supports them, and which claims the resulting simulation can justify.
Four sources, different claims
We distinguish four common construction sources: human-authored archetypes, model-generated personas, population-sampled personas, and trace-grounded personas. Each is useful, but each supports a different validity claim. A designer-authored archetype can make a scenario concrete without representing a population; a trace-grounded persona may closely reflect one person while saying little about everyone else.
A lightweight reporting standard
We propose six minimum items for persona-based simulation studies to report:
- Persona provenance: how the personas were constructed.
- Grounding evidence: the datasets, surveys, traces, or theory used.
- Selection or sampling logic: the target population, sampling frame, weighting, and filtering.
- Internal consistency: checks for contradictory or infeasible attributes.
- Enactment checks: evidence that the model preserves and uses persona attributes.
- Intended use and inference: the claims supported and what remains out of scope.
The goal is not to impose one universal persona method. It is to make construction choices inspectable, comparisons more meaningful, and limitations harder to hide behind plausible prose.