Large language models are increasingly being used to act like people: answering surveys, shopping for products, interacting with assistants, navigating software, and participating in simulated societies.
These systems are often guided by a persona, such as a description of the user's background, preferences, personality, goals, or experience. But a detailed persona does not automatically produce a realistic simulated user. A synthetic customer may sound convincing while behaving far more patiently than a real customer, and a simulated survey respondent may give coherent answers while still misrepresenting the population it is supposed to represent. Our survey argues that simulated users should therefore be judged not by whether they seem generally human-like, but by whether they are fit for the particular purpose for which they were built.
Start with the claim, not the persona
A common approach is to create a persona first and then ask what it can be used for. We propose reversing that process: begin by specifying what the simulation is intended to reproduce. That claim might concern a person's response to a survey question, a sequence of choices in a recommender system, a customer's interaction with a service assistant, behavior inside a social group, or the way a novice navigates a computer interface.
Each claim requires different information about the user. Demographic attributes and attitudes may matter for public-opinion simulation, while a customer-service simulator may instead need the user's hidden goal, incomplete knowledge, patience, and conditions for abandoning the interaction. More detail is not always better; the useful information is the information that actually matters for the behavior being simulated.
A fitness-for-purpose framework
The framework organizes persona-conditioned simulation as a chain of five connected decisions: the simulation claim, the required user state, persona construction, the behavioral mechanism, and the validity evidence. It also includes feedback: interaction history, memory, and changes in the environment may update the simulated user over time, while evaluation may force researchers to narrow—or reject—the claims they originally hoped to make.
A weakness at one stage cannot usually be repaired by strength at another. A detailed persona may be poorly grounded, grounded attributes may have little effect on behavior, and plausible outputs may still fail to match the real process or population of interest.
Personas are not the whole user
We use persona to describe relatively persistent information, such as background, identity, values, skills, personality, social role, and enduring preferences. Many behaviors, however, depend on a broader and more changeable user state that can include a current goal, practical constraints, prior interactions, available knowledge, relationships, memory, trust, frustration, or fatigue.
This distinction becomes especially important in longer interactions. A shopper's general taste may remain stable while their budget, urgency, and reaction to previous recommendations change. A service user may begin cooperatively but become frustrated after repeated errors. A social agent's behavior may depend not only on personality, but also on relationships, norms, and earlier events. A system that claims to simulate behavior over time therefore needs more than a biography; it needs a defensible account of what changes, why it changes, and how those changes affect later behavior.
Where the persona comes from matters
Personas may be written by designers, extracted from interaction histories, reconstructed from documents, sampled from surveys, or generated by another language model. These sources support different kinds of claims. A human-authored profile may be easy to understand and control, but it does not necessarily represent a real population. A trace-derived profile is tied to observed behavior, but may reflect only a narrow or outdated slice of someone's life. Survey-based personas can support population-level analysis when sampling and weighting are explicit. Model-generated personas can be produced at enormous scale, but their apparent diversity may reflect the model's own assumptions rather than the population researchers intend to study.
For this reason, the framework distinguishes between information that was directly observed, sampled from a defined distribution, inferred from other evidence, imputed because something was missing, or generated synthetically. Without those distinctions, invented detail can easily appear more grounded than it really is.
Different applications fail in different ways
There is no single standard for a good simulated user because different applications make different claims. In survey simulation, believable individual answers do not guarantee a representative population; models may flatten variation within groups or reproduce stereotypes. In recommender systems, simulated users may become unnaturally cooperative “preference oracles” who explain their tastes more clearly and consistently than real people.
Customer-service simulators present a different challenge. Real users may withhold information, misunderstand questions, correct the assistant, reject faulty assumptions, lose patience, or abandon the task. A simulator that simply follows the assistant's preferred workflow can make a weak system appear much stronger than it would be in practice.
Social and computer-use simulations introduce still other requirements. Fluent dialogue does not establish a valid social process, and an agent optimized to complete a task is not necessarily a realistic human user. Human behavior includes hesitation, mistakes, help-seeking, recovery, and abandonment, and those patterns should vary with differences in skill, experience, relationships, and context.
Four questions for evaluation
The manuscript distinguishes four complementary forms of validity. Persona fidelity asks whether the model actually preserves and uses the supplied persona. Behavioral and interaction validity asks whether individual actions and longer trajectories resemble the target human process. Task and ecological validity asks whether evaluation with simulated users exposes the same capabilities and failures that would appear with real users under realistic conditions. Distributional validity asks whether a collection of simulated users represents the intended population, including subgroup frequencies, attribute relationships, rare cases, and variation within groups.
Passing one form of validity does not guarantee the others. A simulator may follow its persona faithfully even when the persona itself is poorly grounded. It may reproduce population averages while generating implausible individual behavior, or produce convincing conversations without exposing the same system failures that human users would reveal.
Simulated users should complement human evidence
Persona-conditioned simulators can help pretest questionnaires, generate difficult interaction scenarios, uncover system failures, explore hypotheses, and expand the range of users and conditions considered during evaluation. Their value is real, but it depends on demonstrated correspondence with the use for which they are intended.
They should not silently replace human participants or affected communities. High-impact claims about populations, products, policies, or deployed systems still require calibration against human evidence and a clear account of the remaining gap between simulation and reality. The central recommendation is therefore straightforward: make the intended claim explicit, represent the user state needed for that claim, document where the information came from, explain how it affects behavior, and evaluate the simulator against the human process it is supposed to represent.
The future of persona-based simulation depends not only on building larger persona libraries or more expressive agents, but on knowing what those agents are actually valid for.