Language model systems are increasingly deployed to stand in for human users. Recommender systems simulate consumer cohorts, social scientists deploy synthetic respondents to test survey instruments, and product teams use simulated users to test software flows and conversational bots prior to release. In each of these settings, an identity profile is assigned in a system prompt, and downstream decisions depend on the model acting faithfully on that identity without constant reminders.
Yet existing evaluations focus almost exclusively on conversational role-play: maintaining a stylistic tone, recalling biographical trivia, or avoiding out-of-character persona drift. These criteria only test the conversational surface. When an agent acts on a user's behalf, the decisive question is behavioral: does a latent trait observably govern what the agent decides in an everyday situation where that trait is never explicitly mentioned?
Consider the gap between prose and action: a vegan persona ordering dinner must not add meat to the cart; a risk-averse persona rebalancing a portfolio must avoid speculative assets. If a persona merely flavors the agent's conversational style while the underlying decision defaults to generic LLM tendencies, the simulation provides no predictive value. To systematically evaluate authentic behavioral steerability, we introduce AgentPersonaBench (APB).
1. Benchmark Architecture and Design Principles
Evaluating behavioral steerability requires isolating individual traits without reducing the task to trivial instruction following. APB is built around four core methodological commitments:
Latent conditioning over explicit prompting
Standard benchmarks often evaluate identity with isolated, one-line prompts (such as "You are a vegan. Order a meal."). In that setup, the model merely performs simple instruction following on a single salient rule. In production deployments, however, agents operate under persistent user profiles with numerous background details.
APB embeds each target trait within a complete synthetic persona profile sampled from a one-million persona corpus. The profile incorporates realistic background context—including demographic foundations, lifestyle habits, values, and communication styles. The target trait is unannounced and must actively compete for attention against surrounding context, reproducing the distracting conditions of real-world deployments.
Unprompted operational errands
Tasks are formulated as ordinary everyday errands—such as ordering dinner from a menu, choosing a study spot, or completing an onboarding workflow. The task instruction never mentions the tested trait, never provides hints, and never announces that an evaluation is taking place. Ground-truth conditions remain strictly isolated from the agent's view.
Four interaction surfaces of increasing realism
Persona adherence is evaluated across four distinct modalities spanning 2,460 tasks:
- Survey (502 tasks): Structured questionnaires with randomized option ordering and counterbalanced answer positions to eliminate selection biases.
- Chat (565 tasks): Multi-turn conversations where a scripted interlocutor applies conversational pressure, actively nudging the agent toward off-persona choices (for example, suggesting non-vegan dishes or encouraging speculative financial risks).
- Web (707 tasks): Interactive browser environments where decisions are recorded via DOM state, shopping carts, and interactive web elements.
- App (686 tasks): Desktop GUI software environments where agents interact with native applications and write state files to disk.
Deterministic artifact verification
APB verifies model performance strictly from observable environment artifacts rather than conversational self-reports or post-hoc justifications. Approximately four-fifths (78.5%) of all checks are deterministic code predicates evaluated over structured artifacts: selected option IDs, DOM tree states, and application output files. A standardized LLM rubric judge is reserved solely for free-text dialogue turns in chat tasks.
Every task admitted to APB underwent automated audits and human review to eliminate guessing shortcuts, option position biases, and superficial prompt leaks before inclusion in the final benchmark suite.
2. The Leaderboard
We evaluated 20 frontier model arms from OpenAI, Anthropic, Google, xAI, Zhipu AI, Qwen, Moonshot, and DeepSeek under standardized execution harnesses at medium reasoning effort. Performance is reported as the full-pass rate—the percentage of tasks on which every requirement was satisfied. Execution timeouts and environment errors are excluded rather than scored as behavioral failures.
| # | Model | Overall | Survey | Chat | Web | App | Single | Multi | $/task |
|---|---|---|---|---|---|---|---|---|---|
| 1 | gemini-3-8-flashGoogle | 84.7 | 88.3 | 79.6 | 85.0 | 86.0 | 86.7 | 82.8 | $0.092 |
| 2 | opus-5-5Anthropic | 82.9 | 84.3 | 79.9 | 85.3 | 81.8 | 85.0 | 80.9 | $0.416 |
| 3 | gemini-3-7-flashGoogle | 82.7 | 86.1 | 77.7 | 83.5 | 83.7 | 85.0 | 80.5 | $0.079 |
| 4 | gpt-6-astraOpenAI | 82.1 | 83.3 | 80.7 | 86.1 | 78.1 | 83.6 | 80.6 | $0.869 |
| 5 | gpt-6-solOpenAI | 76.0 | 86.7 | 71.3 | 78.6 | 69.3 | 82.6 | 69.6 | $0.170 |
| 6 | opus-5Anthropic | 74.9 | 85.1 | 61.7 | 76.7 | 76.2 | 85.4 | 64.7 | $0.510 |
| 7 | opus-4-8Anthropic | 74.0 | 83.1 | 65.2 | 72.0 | 76.5 | 84.4 | 63.9 | $0.478 |
| 8 | gpt-5-6-solOpenAI | 71.4 | 80.7 | 68.2 | 72.0 | 66.3 | 80.8 | 61.8 | $0.323 |
| 9 | kimi-k3Moonshot | 70.1 | 79.0 | 53.7 | 75.0 | 72.2 | 78.1 | 62.4 | $0.331 |
| 10 | grok-4-6xAI | 68.8 | 81.7 | 52.9 | 70.7 | 69.9 | 79.1 | 58.6 | $0.282 |
| 11 | deepseek-v4-pro-0813DeepSeek | 68.0 | 85.3 | 51.9 | 68.6 | N/A | 76.6 | 59.6 | $0.056 |
| 12 | qwen-3-8-flashQwen | 66.8 | 72.1 | 54.6 | 66.9 | 72.8 | 75.3 | 58.6 | $0.015 |
| 13 | glm-5-3GLM | 66.3 | 79.3 | 62.7 | 60.0 | N/A | 81.1 | 51.7 | $0.137 |
| 14 | gpt-5-6-terraOpenAI | 64.1 | 71.5 | 60.1 | 65.2 | 60.8 | 76.8 | 51.8 | $0.152 |
| 15 | glm-5-3-flashGLM | 62.6 | 69.1 | 48.2 | 65.6 | 66.5 | 73.0 | 52.6 | $0.014 |
| 16 | qwen-3-8-27bQwen | 62.6 | 71.1 | 47.8 | 68.0 | 63.0 | 72.4 | 52.9 | $0.038 |
| 17 | deepseek-v4-proDeepSeek | 60.9 | 75.7 | 46.7 | 61.6 | N/A | 69.5 | 52.4 | $0.052 |
| 18 | gpt-6-lunaOpenAI | 59.8 | 68.3 | 60.3 | 57.7 | 55.2 | 74.4 | 45.7 | $0.008 |
| 19 | gpt-5-6-lunaOpenAI | 56.2 | 65.7 | 46.9 | 56.0 | 57.0 | 72.1 | 40.8 | $0.016 |
| 20 | deepseek-v4-flashDeepSeek | 50.5 | 68.7 | 42.0 | 44.3 | N/A | 64.8 | 36.4 | $0.020 |
Four frontier models form a competitive top tier within three points of each other: gemini-3-8-flash (84.7%), opus-5-5 (82.9%), gemini-3-7-flash (82.7%), and gpt-6-astra (82.1%). Furthermore, empirical fidelity does not strictly correlate with inference pricing: Gemini Flash arms achieve top-3 performance at under ten cents per task, while larger flagship arms incur substantially higher inference costs for comparable adherence.
3. Empirical Findings and Diagnostic Insights
Analysis across the 20 model arms and matched task subsets reveals five key behavioral findings with direct consequences for practitioner deployments:
1. High-fidelity persona simulation is achievable
Under unprompted, latent conditioning, top frontier models successfully act on persona traits in more than four out of five tasks. Nine of the 20 evaluated arms exceed 70% full-pass adherence. This confirms that modern language models can reliably sustain implicit identities across diverse operational environments without explicit instruction prompting.
2. Conversational pushback degrades adherence
Multi-turn chat is the lowest-scoring surface for 15 out of 20 model arms. When a conversational assistant applies pushback (for instance, arguing against a stated dietary restriction or encouraging speculative investments), weaker models frequently capitulate.
This vulnerability scales sharply with model capacity: on matched scenarios, gpt-6-astra maintains identical scores across survey and chat (85.0% vs. 85.7%), whereas gpt-5-6-luna drops by 15.7 percentage points under dialogue pressure. For chatbot evaluation studies, model steerability under social pressure is the primary bottleneck.
3. Multi-attribute co-exposure reveals capability cliffs
Real-world user simulation requires agents to embody multiple co-occurring preferences simultaneously. In multi-attribute tasks where three traits are evaluated jointly, frontier models remain resilient (gemini-3-8-flash 82.8%, opus-5-5 80.9%, gpt-6-astra 80.6%).
In contrast, smaller and mid-tier models exhibit sharp capability drops: gpt-5-6-luna falls from 72.1% on single-trait tasks to 40.8% on multi-trait tasks. Notably, 71% to 88% of failed multi-trait tasks still uphold at least one trait, indicating partial adherence rather than complete identity collapse.
4. Cross-surface fragility
To assess cross-surface transfer, we analyzed a matched cohort of 140 scenarios implemented identically across all four media. While gpt-6-astra achieves 80.0% to 89.3% adherence on individual surfaces, its all-surface pass rate falls to 64.3%. For gpt-5-6-luna, all-surface adherence drops to 37.9%.
Furthermore, desktop app tasks exhibit the highest conditional failure rate across every model arm: among scenarios passing surveys, 14.3% fail in desktop environments for Astra, and 30.8% fail for Luna. Practitioners cannot treat survey benchmarks as proxies for interactive web or desktop application behavior.
5. Cross-vendor complementarity and statistical stability
On a common cohort of 1,086 chat and web tasks scored across 18 model arms, models display remarkable behavioral complementarity: Gemini 3.8 successfully passes 53.6% of the tasks failed by GPT-6 Astra. Evaluating critical product decisions across multiple distinct model families provides orthogonal error detection.
Finally, repeat evaluations over three complete runs on the GPT-6 model pair demonstrate that performance differences are statistically robust: the standard deviation across runs is ≤ 2.0 percentage points, confirming that APB scores reflect durable architectural capabilities rather than sampling noise.
4. Practical Takeaways for Simulated User Deployment
These findings point to three concrete recommendations for teams building synthetic user panels and evaluating interactive software:
- Validate on the target deployment surface: Never assume that a model passing survey-based evaluations will behave faithfully within interactive web forms or native desktop workflows. Benchmark specifically on the intended interaction modality.
- Prioritize frontier reasoning for conversational and multi-faceted personas: When simulated users must converse with assistants or balance multiple personal habits, weaker models suffer steep fidelity drop-offs. High-reasoning models are essential for conversational resistance.
- Cross-verify with diverse model families: Because frontier models from different providers exhibit complementary failure modes, running synthetic panels with personas enacted across two independent model families provides the highest confidence in simulated user research.
5. Open-Source Resources and Benchmark Access
All benchmark artifacts, environments, and evaluation tooling are fully open-source and publicly accessible:
- arXiv Research Paper: Formal mathematical formulations, related work, task admission audits, and extended analysis.
- AgentPersonaBench GitHub Repository: All 2,460 tasks with SHA-256 integrity manifests, standardized evaluation harnesses for all 20 model arms, and leaderboard reproduction scripts.
- MatrAIx Persona 1M on Hugging Face: The complete synthetic persona population corpus.
- MatrAIx Interactive Playground: Interactive environment for deploying and testing simulated persona cohorts against custom applications.
- Simulating the World with 8.3 Billion Personas: The technical report detailing the underlying persona generation architecture.
Responsible Use: APB personas are synthetic instruments designed for empirical software evaluation and behavioral testing, not representations of real identifiable individuals. Simulated users provide rapid early-stage feedback across diverse customer segments, but critical decisions should continue to be validated alongside real human user research.
6. Citation
If you use AgentPersonaBench or build upon our evaluation findings, please cite our publication:
@misc{huang2026agentpersonabenchbenchmarkingpersonadrivenuser,
title={AgentPersonaBench: Benchmarking Persona-Driven User Simulation},
author={Jintao Huang and Yifan Wang and Hongyu Shen and Yi Daniel Lu and Shirley Huang and Minsik Oh and Yewen Wang and Muhammad Ahmed Mohsin and Zhen Xu and Yilan Fan and Zichen Yuan and Ahsan Bilal and Zibu Wei and Sankalp Jajee and Henry Gagnier and Saksham Kapoor and Jicheng Wang and Qianfeng Wen and Yixuan He and Steven Dillmann and Jiashu He and Yucheng Lu and Linqiang Guo and Danyang Zhang and Shi Bo and Raunak Mondal and Haixiang Tang and Weihang Xiao and Allen Nie and Jing Tang and Yueying Li and Yifan Simon Liu and Jianheng Hou and Dianzhuo Wang and Qianyu Zhu and Zhixu Silvia Tao and Zhejian Peng and Zihan Wang and Ishan Gupta and Jinxuan Fan and Wanting Jiang and Shushu Liang and Chenxi Qiu and Yijun Wang and Xiaomin Li and Yuexing Hao},
year={2026},
eprint={2610.04379},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.04379},
}
Evidence note. All reported metrics are directly reproduced from the AgentPersonaBench paper and the verified leaderboard dataset: 2,460 tasks, 867 traits, 20 model arms, cross-surface matched cohorts, multi-attribute evaluations, and three-repeat stability trials. Per-task evaluation traces and recomputation scripts are available in the public repository.
