Home/Research/Publication

PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications

A plug-and-play framework that connects persona-driven users to surveys, chatbots, and web applications through one evaluation workflow.

Accepted COLM 2026 @ SocialSim
MatrAIx Research Community
July 26, 2026
3 min read
PersonaEval dashboards for survey, chatbot, and web application evaluations
PersonaEval runs the same persona-driven evaluation workflow across multiple application interfaces.

Real user studies remain essential, but recruiting participants and repeating studies after every product change is expensive and slow. PersonaEval explores a complementary layer: simulated users who can interact with different applications in repeatable, parallel runs.

Our system paper, accepted to the COLM 2026 Workshop on Social Simulation with LLMs, introduces a modular workflow that separates the simulated user from the application under test. A Persona API turns an existing persona profile into user behavior, while a Task API packages the interface, interaction protocol, environment, and evaluation form.

One workflow, three interaction settings

Because the task adapter can change without rebuilding the persona side, the same population can evaluate different systems while interaction trajectories and outcomes stay in a common format.

PersonaEval treats simulated-user evaluation as infrastructure: choose a persona population, connect an application adapter, run interactions concurrently, and aggregate both outcomes and traces.

What the first results show

We evaluated six application entries with 50 relevant Nemotron personas each. The results show application-level differences and meaningful variation across persona groups. Multi-turn chatbots produced wider outcome spreads as users revealed constraints and negotiated responses, while survey and web tasks produced more concentrated scores.

Human-judged persona alignment remained high across the three settings. Representative traces also show personas shaping goals, product constraints, critiques, and final ratings. These results are an initial demonstration, not a substitute for human calibration; future work will compare simulated outcomes directly with real-user data and expand to more application domains.