Home/Research/Publication

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

When language models act as simulated users, do latent persona profiles actually govern their downstream actions? We evaluate 20 frontier models across 2,460 everyday tasks spanning surveys, multi-turn dialogues, web interfaces, and desktop applications.

MatrAIx
Jintao Huang, Yuexing Hao, Xiaomin Li, and Core Team
October 5, 2026
10 min read

Language model systems are increasingly deployed to stand in for human users. Recommender systems simulate consumer cohorts, social scientists deploy synthetic respondents to test survey instruments, and product teams use simulated users to test software flows and conversational bots prior to release. In each of these settings, an identity profile is assigned in a system prompt, and downstream decisions depend on the model acting faithfully on that identity without constant reminders.

Yet existing evaluations focus almost exclusively on conversational role-play: maintaining a stylistic tone, recalling biographical trivia, or avoiding out-of-character persona drift. These criteria only test the conversational surface. When an agent acts on a user's behalf, the decisive question is behavioral: does a latent trait observably govern what the agent decides in an everyday situation where that trait is never explicitly mentioned?

A persona card with preferences, habits, values, and style feeds three everyday situations: choosing a place to read, replying to a friend, and buying a desk lamp. The observed choice either holds or violates the persona's preference.
Does a persona govern downstream behavior? Describing an individual profile is not enough: latent traits must dictate actual choices—such as selecting a quiet reading nook, keeping a reply concise, or placing the budget-friendly lamp in the cart.

Consider the gap between prose and action: a vegan persona ordering dinner must not add meat to the cart; a risk-averse persona rebalancing a portfolio must avoid speculative assets. If a persona merely flavors the agent's conversational style while the underlying decision defaults to generic LLM tendencies, the simulation provides no predictive value. To systematically evaluate authentic behavioral steerability, we introduce AgentPersonaBench (APB).

2,460audited evaluation tasks
867distinct personal traits
4interaction surfaces
20frontier model arms evaluated

1. Benchmark Architecture and Design Principles

Evaluating behavioral steerability requires isolating individual traits without reducing the task to trivial instruction following. APB is built around four core methodological commitments:

Latent conditioning over explicit prompting

Standard benchmarks often evaluate identity with isolated, one-line prompts (such as "You are a vegan. Order a meal."). In that setup, the model merely performs simple instruction following on a single salient rule. In production deployments, however, agents operate under persistent user profiles with numerous background details.

APB embeds each target trait within a complete synthetic persona profile sampled from a one-million persona corpus. The profile incorporates realistic background context—including demographic foundations, lifestyle habits, values, and communication styles. The target trait is unannounced and must actively compete for attention against surrounding context, reproducing the distracting conditions of real-world deployments.

Unprompted operational errands

Tasks are formulated as ordinary everyday errands—such as ordering dinner from a menu, choosing a study spot, or completing an onboarding workflow. The task instruction never mentions the tested trait, never provides hints, and never announces that an evaluation is taking place. Ground-truth conditions remain strictly isolated from the agent's view.

Four interaction surfaces of increasing realism

Persona adherence is evaluated across four distinct modalities spanning 2,460 tasks:

Top row: a vegan persona chooses meals across survey, chat, web, and app, and gets a held or violated verdict. Bottom row: a developer persona with three pinned coding habits writes the same function on all four surfaces and gets a score from zero to three.
Multi-surface task execution. (a) Single-attribute tasks isolate one latent trait (e.g., dietary constraint) across all four interaction surfaces. (b) Multi-attribute tasks evaluate whether multiple co-occurring habits can be sustained simultaneously.

Deterministic artifact verification

APB verifies model performance strictly from observable environment artifacts rather than conversational self-reports or post-hoc justifications. Approximately four-fifths (78.5%) of all checks are deterministic code predicates evaluated over structured artifacts: selected option IDs, DOM tree states, and application output files. A standardized LLM rubric judge is reserved solely for free-text dialogue turns in chat tasks.

Three panels: sample the persona from the one-million persona corpus with a pinned trait, build a survey, chat, web, or app environment that exposes the choice, and design a deterministic verifier or a fixed LLM judge that defines what counts as held.
Task construction pipeline. Target traits are paired with contrast values, embedded into complete persona profiles, staged across four interaction media, and evaluated via deterministic predicates.

Every task admitted to APB underwent automated audits and human review to eliminate guessing shortcuts, option position biases, and superficial prompt leaks before inclusion in the final benchmark suite.

2. The Leaderboard

We evaluated 20 frontier model arms from OpenAI, Anthropic, Google, xAI, Zhipu AI, Qwen, Moonshot, and DeepSeek under standardized execution harnesses at medium reasoning effort. Performance is reported as the full-pass rate—the percentage of tasks on which every requirement was satisfied. Execution timeouts and environment errors are excluded rather than scored as behavioral failures.

#ModelOverallSurveyChatWebAppSingleMulti$/task
1gemini-3-8-flashGoogle84.788.379.685.086.086.782.8$0.092
2opus-5-5Anthropic82.984.379.985.381.885.080.9$0.416
3gemini-3-7-flashGoogle82.786.177.783.583.785.080.5$0.079
4gpt-6-astraOpenAI82.183.380.786.178.183.680.6$0.869
5gpt-6-solOpenAI76.086.771.378.669.382.669.6$0.170
6opus-5Anthropic74.985.161.776.776.285.464.7$0.510
7opus-4-8Anthropic74.083.165.272.076.584.463.9$0.478
8gpt-5-6-solOpenAI71.480.768.272.066.380.861.8$0.323
9kimi-k3Moonshot70.179.053.775.072.278.162.4$0.331
10grok-4-6xAI68.881.752.970.769.979.158.6$0.282
11deepseek-v4-pro-0813DeepSeek68.085.351.968.6N/A76.659.6$0.056
12qwen-3-8-flashQwen66.872.154.666.972.875.358.6$0.015
13glm-5-3GLM66.379.362.760.0N/A81.151.7$0.137
14gpt-5-6-terraOpenAI64.171.560.165.260.876.851.8$0.152
15glm-5-3-flashGLM62.669.148.265.666.573.052.6$0.014
16qwen-3-8-27bQwen62.671.147.868.063.072.452.9$0.038
17deepseek-v4-proDeepSeek60.975.746.761.6N/A69.552.4$0.052
18gpt-6-lunaOpenAI59.868.360.357.755.274.445.7$0.008
19gpt-5-6-lunaOpenAI56.265.746.956.057.072.140.8$0.016
20deepseek-v4-flashDeepSeek50.568.742.044.3N/A64.836.4$0.020
Full-pass adherence rate (%) across interaction surfaces, single- and multi-attribute tasks, and overall. N/A marks models evaluated without native desktop GUI tool execution. Cost reflects mean inference cost per task.

Four frontier models form a competitive top tier within three points of each other: gemini-3-8-flash (84.7%), opus-5-5 (82.9%), gemini-3-7-flash (82.7%), and gpt-6-astra (82.1%). Furthermore, empirical fidelity does not strictly correlate with inference pricing: Gemini Flash arms achieve top-3 performance at under ten cents per task, while larger flagship arms incur substantially higher inference costs for comparable adherence.

3. Empirical Findings and Diagnostic Insights

Analysis across the 20 model arms and matched task subsets reveals five key behavioral findings with direct consequences for practitioner deployments:

1. High-fidelity persona simulation is achievable

Under unprompted, latent conditioning, top frontier models successfully act on persona traits in more than four out of five tasks. Nine of the 20 evaluated arms exceed 70% full-pass adherence. This confirms that modern language models can reliably sustain implicit identities across diverse operational environments without explicit instruction prompting.

2. Conversational pushback degrades adherence

Multi-turn chat is the lowest-scoring surface for 15 out of 20 model arms. When a conversational assistant applies pushback (for instance, arguing against a stated dietary restriction or encouraging speculative investments), weaker models frequently capitulate.

This vulnerability scales sharply with model capacity: on matched scenarios, gpt-6-astra maintains identical scores across survey and chat (85.0% vs. 85.7%), whereas gpt-5-6-luna drops by 15.7 percentage points under dialogue pressure. For chatbot evaluation studies, model steerability under social pressure is the primary bottleneck.

3. Multi-attribute co-exposure reveals capability cliffs

Real-world user simulation requires agents to embody multiple co-occurring preferences simultaneously. In multi-attribute tasks where three traits are evaluated jointly, frontier models remain resilient (gemini-3-8-flash 82.8%, opus-5-5 80.9%, gpt-6-astra 80.6%).

In contrast, smaller and mid-tier models exhibit sharp capability drops: gpt-5-6-luna falls from 72.1% on single-trait tasks to 40.8% on multi-trait tasks. Notably, 71% to 88% of failed multi-trait tasks still uphold at least one trait, indicating partial adherence rather than complete identity collapse.

A persona that passes a static questionnaire can readily fail the identical errand in an interactive web browser or desktop application. High score on one interaction surface does not certify adherence across others.

4. Cross-surface fragility

To assess cross-surface transfer, we analyzed a matched cohort of 140 scenarios implemented identically across all four media. While gpt-6-astra achieves 80.0% to 89.3% adherence on individual surfaces, its all-surface pass rate falls to 64.3%. For gpt-5-6-luna, all-surface adherence drops to 37.9%.

Furthermore, desktop app tasks exhibit the highest conditional failure rate across every model arm: among scenarios passing surveys, 14.3% fail in desktop environments for Astra, and 30.8% fail for Luna. Practitioners cannot treat survey benchmarks as proxies for interactive web or desktop application behavior.

5. Cross-vendor complementarity and statistical stability

On a common cohort of 1,086 chat and web tasks scored across 18 model arms, models display remarkable behavioral complementarity: Gemini 3.8 successfully passes 53.6% of the tasks failed by GPT-6 Astra. Evaluating critical product decisions across multiple distinct model families provides orthogonal error detection.

Finally, repeat evaluations over three complete runs on the GPT-6 model pair demonstrate that performance differences are statistically robust: the standard deviation across runs is ≤ 2.0 percentage points, confirming that APB scores reflect durable architectural capabilities rather than sampling noise.

4. Practical Takeaways for Simulated User Deployment

These findings point to three concrete recommendations for teams building synthetic user panels and evaluating interactive software:

5. Open-Source Resources and Benchmark Access

All benchmark artifacts, environments, and evaluation tooling are fully open-source and publicly accessible:

Responsible Use: APB personas are synthetic instruments designed for empirical software evaluation and behavioral testing, not representations of real identifiable individuals. Simulated users provide rapid early-stage feedback across diverse customer segments, but critical decisions should continue to be validated alongside real human user research.

6. Citation

If you use AgentPersonaBench or build upon our evaluation findings, please cite our publication:

@misc{huang2026agentpersonabenchbenchmarkingpersonadrivenuser,
      title={AgentPersonaBench: Benchmarking Persona-Driven User Simulation}, 
      author={Jintao Huang and Yifan Wang and Hongyu Shen and Yi Daniel Lu and Shirley Huang and Minsik Oh and Yewen Wang and Muhammad Ahmed Mohsin and Zhen Xu and Yilan Fan and Zichen Yuan and Ahsan Bilal and Zibu Wei and Sankalp Jajee and Henry Gagnier and Saksham Kapoor and Jicheng Wang and Qianfeng Wen and Yixuan He and Steven Dillmann and Jiashu He and Yucheng Lu and Linqiang Guo and Danyang Zhang and Shi Bo and Raunak Mondal and Haixiang Tang and Weihang Xiao and Allen Nie and Jing Tang and Yueying Li and Yifan Simon Liu and Jianheng Hou and Dianzhuo Wang and Qianyu Zhu and Zhixu Silvia Tao and Zhejian Peng and Zihan Wang and Ishan Gupta and Jinxuan Fan and Wanting Jiang and Shushu Liang and Chenxi Qiu and Yijun Wang and Xiaomin Li and Yuexing Hao},
      year={2026},
      eprint={2610.04379},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.04379}, 
}

Evidence note. All reported metrics are directly reproduced from the AgentPersonaBench paper and the verified leaderboard dataset: 2,460 tasks, 867 traits, 20 model arms, cross-surface matched cohorts, multi-attribute evaluations, and three-repeat stability trials. Per-task evaluation traces and recomputation scripts are available in the public repository.