A real person is expensive for a good reason. They must be found, screened, compensated, scheduled, supported, and treated ethically. Their time is scarce. A simulated user can be copied, reset, run in parallel, and asked to repeat the same task after every product change.
That difference changes the economics of evaluation. It does not make the simulated user a drop-in replacement for a person. It creates a much cheaper layer of evidence that can run before and between human studies.
MatrAIx is built to make that layer practical at scale: large populations of persona-conditioned agents that probe a product the way a varied user base would, then hand the sharpest questions to real people. This article puts numbers on the underlying economics using sample sizes appropriate to different evaluation modes: 10,000 survey responses, 1,000 long conversations, 500 product journeys, 100,000 annotated outputs, and a 28,000-visitor A/B test. We compare transparent public prices for participant recruitment with token-based inference costs. The estimates are not vendor quotes or results from a completed MatrAIx benchmark; they are reproducible scenarios whose assumptions can be replaced with a team's own traffic, model, token budget, labor rate, and confidence requirements.
| Evaluation | Human collection | Human time | Simulation pass | Simulation time | Order of magnitude |
|---|---|---|---|---|---|
| 10K concept surveys | $28,560+ | 3 to 10 days | $3.60 to $104two passes | 15 to 45 min | ~270 to 7,900× |
| 1K support chats | $8,568+ | 1 to 3 weeks | $3 to $87 | 30 to 90 min | ~100 to 2,900× |
| 500 product journeys | $2,142+ | 1 to 2 weeks | $1.44 to $42plus tool cost | 1 to 3 hours | ~50 to 1,500× |
| 100K outputs, 3 labels each | $85,680 | 2 to 4 weeks | $12.60 to $365 | under 1 hour | ~230 to 6,800× |
| A/B test, 3% → 3.6% | production traffic | 7 days or more | $3.60 to $104 | 15 to 60 min | triage, not proof |
Variable data-collection cost only, before shared research labor. Human figures use Prolific fair-pay rates; simulation figures span an economical to premium open model on public token prices. The A/B row compares different kinds of evidence, not equivalent line items. Full assumptions and sources are below.
Start with the units, not the headline
Human research and AI simulation meter different things. Human platforms charge for participant time, recruitment, targeting, scheduling, and incentives. Model providers charge for input and output tokens. Product teams also pay for study design, instrumentation, analysis, and decision-making in both cases.
participants × task hours × reward rate × (1 + platform fee)
MODEL INFERENCE
(input tokens ÷ 1M × input price) + (output tokens ÷ 1M × output price)
For a public reference point, Prolific recommends at least $12 per participant-hour and normally adds a 42.8% corporate platform fee (33.3% for academic and nonprofit customers). Its minimum allowed rate is $8 per hour.1 User Interviews lists recruiting at $49 per completed session, or $98 with advanced B2B targeting, and explicitly excludes participant incentives.2 Respondent lists $40 per on-demand session or $34 in a prepaid 63-session bundle.3 At 10,000 participants, specialist recruiting services would normally require negotiated volume terms, so our large-N examples use Prolific’s transparent fair-pay formula rather than multiplying a boutique recruiting fee.
Those per-session numbers are only the recruiting line. A finished study costs much more. Drive Research, a United States market-research firm, publishes ballpark budgets of $5,000 to $15,000 for a 400-response online survey, $7,000 to $20,000 or more per focus group, and $20,000 to $50,000 for an in-person survey of the same size.9 Those are the real numbers a simulation pass is measured against.
For model inference, we show three public Together AI tiers rather than pretend there is one “AI cost”: economical Gemma 3n E4B at $0.06/M input and $0.12/M output tokens, Qwen3.5 9B at $0.17/M input and $0.25/M output, and premium DeepSeek V4 Pro at $1.74/M input and $3.48/M output.4 These are list-price examples, not quality endorsements. Actual enterprise rates may differ, and batching, prompt caching, self-hosting, or volume discounts can lower the bill. A small model may be suitable for structured surveys or labels but fail at nuanced behavior; a premium model can cost more and still be poorly calibrated.
It is fair to object that AI is no longer cheap either. A per-token price can look intimidating in isolation. The number that matters is the cost of a whole task, and it stays small: a full 10,000-persona survey pass on the economical model runs about $1.80, roughly the price of a coffee. That price is also falling fast. Epoch AI finds that inference prices have dropped rapidly across tasks, and the Stanford AI Index reports that the cost of running a system at GPT-3.5-level quality fell more than 280-fold between November 2022 and October 2024.10 Inference is the rare input that keeps getting cheaper. Human time does not.
They price participant time and the published platform fee. They do not include researcher salaries, screener design, recruitment operations, consent handling, no-shows, facilitation, transcription, qualitative coding, analytics, procurement, or the opportunity cost of waiting. The AI side likewise excludes shared engineering and calibration, but once an evaluation harness exists, its marginal cost is dominated by inference and tool execution.
Five comparisons
Every comparison isolates the variable cost of collecting evaluation data. Shared product setup and analyst time are excluded because both approaches need them, dollar amounts are rounded, and the token budgets are illustrative, not measured MatrAIx workloads.
On timing, human estimates include realistic recruiting and fielding windows, not just paid participant-hours, while simulation estimates assume an implemented workflow and workload-appropriate parallelism. Both are planning envelopes rather than benchmarks: tight rate limits can stretch hours into days, and batching can pull them back.
Ten thousand reactions to a ten-minute concept survey
Human collection. At Prolific’s recommended $12/hour, ten minutes costs $2 per respondent. Ten thousand responses cost $20,000 in rewards plus $8,560 in platform fees: $28,560, before survey design, quota management, cleaning, and analysis.
Simulation. Give each persona 2,000 input tokens of concept, persona, and instructions and allow 500 output tokens. Ten thousand runs consume 20M input and 5M output tokens: $1.80 on the economical model, $4.65 on Qwen3.5 9B, or $52.20 on the premium example. Running an equally sized second pass for scoring or replication makes that $3.60–$104.40.
A 10% human calibration sample means 1,000 real respondents and adds about $2,856, producing a hybrid variable total near $2,860–$2,960. That is still roughly 90% below full human collection, but it validates only the measured dimensions and sampled population.
One thousand users, each in a thirty-minute support conversation
Human collection. One thousand half-hour sessions require 500 participant-hours. At $12/hour plus the 42.8% fee, variable collection cost is $8,568 before moderation, transcripts, coding, or analysis. Specialist recruiting would cost materially more.
Simulation. A generous budget of 15,000 cumulative input tokens and 5,000 output tokens per multi-turn session totals 15M input and 5M output tokens: $1.50 on the economical model, $3.80 on Qwen3.5 9B, or $43.50 on DeepSeek V4 Pro. Doubling the run for an independent judge or replication yields $3–$87.
This is where simulation is naturally strong: reproducible multi-turn pressure tests across language, patience, expertise, and intent. It is naturally weak at proving that real customers will trust, understand, or return to the product.
Five hundred users across five product-journey segments
Qualitative discovery can work with small groups (Nielsen Norman Group recommends five participants for a typical single-group qualitative study), but quantitative coverage and subgroup comparison require larger samples.5 Here we compare 500 unmoderated 15-minute journeys: 100 observations per segment.
Human collection. At $12/hour, a 15-minute task pays $3. Five hundred completions cost $1,500 in rewards plus $642 in platform fees: $2,142 before study operations and analysis.
Simulation. Five hundred persona-conditioned journeys, budgeted generously at 40,000 input and 4,000 output tokens each, consume 20M input and 2M output tokens: $1.44 on the economical model, $3.90 on Qwen3.5 9B, or $41.76 on the premium example. Browser execution, screenshots, vision input, and sandbox time are additional and can exceed the text-model bill.
Five hundred synthetic journeys are not equivalent to 500 observed people. They provide broad automated coverage. Use them to locate brittle flows and subgroup hypotheses; use people to reveal perception, accessibility, workarounds, emotion, and problems the simulator was never prompted to imagine.
One hundred thousand outputs, three independent labels each
Human collection. If one label takes one minute, 300,000 labels consume 5,000 hours. At $12/hour plus Prolific’s fee, collection costs $85,680. Specialist medical, legal, safety, or coding judgments should pay more, so this is a floor for general participants, not a quote for expert annotation.
AI evaluation. Three independent model passes per item, each using 500 input and 100 output tokens, consume 150M input and 30M output tokens: $12.60 on the economical model, $33 on Qwen3.5 9B, or $365.40 on the premium example. A human audit of 10% of items at one minute each adds about $2,856, yielding a hybrid variable cost of roughly $2,869–$3,221.
The audit must be designed around risk, not convenience. Random checks estimate average error; targeted checks should cover rare classes, model disagreements, protected groups, high-consequence outputs, and distribution shifts.
Finding a 20% relative lift from a 3% baseline
This comparison is different. Ordinary production visitors are usually not paid, so the primary cost is traffic, elapsed time, instrumentation, exposure to a weaker variant, and analyst attention, not a participant invoice.
Using a conventional two-proportion approximation with 95% significance, 80% power, a 3.0% baseline, and a minimum detectable lift to 3.6%, the test needs roughly 14,000 eligible visitors per arm, or about 28,000 total. At 4,000 eligible visitors per day, collection takes about seven days. Optimizely likewise defines duration as total visitors divided by average daily visitors and recommends at least one seven-day business cycle.6
Simulation. Ten thousand pre-launch journeys at 4,000 input and 1,000 output tokens each cost about $3.60–$104.40 across our three model tiers. They can identify obvious regressions, compare behavior under controlled personas, and eliminate weak variants before production.
Simulation cannot establish real conversion lift. A synthetic user is generated by a model, not randomly assigned from the production population. It can reduce how many variants reach the expensive experiment; it cannot turn a modeled preference into causal evidence.
What the dramatic ratio leaves out
The token bill is often the smallest line item. A serious simulator needs persona construction, an environment adapter, task design, observability, failure recovery, scoring, calibration, and human review. If those fixed costs are charged to one tiny study, simulation can be the more expensive approach. Its economics improve when the same infrastructure is reused after every release.
Environment integration, persona design, rubrics, telemetry, calibration, and engineering.
Model tokens, browser or tool execution, retries, judge passes, storage, and review.
Recruitment, incentives, facilitation, consent, scheduling, analysis, and expert adjudication.
The business or social cost of believing a cheap but invalid simulated result.
The last category dominates in healthcare, finance, employment, accessibility, safety, and other high-consequence settings. Saving $8,000 on evaluation is not a saving if the simulator systematically erases a vulnerable group or approves a harmful design.
Models can evaluate models, but they inherit their biases
There is real evidence that strong language models can approximate some human preferences. The MT-Bench and Chatbot Arena study reported over 80% agreement between strong LLM judges and human preferences, comparable to inter-human agreement in its setting. The same paper documents position, verbosity, self-enhancement, and reasoning biases.7 G-Eval found stronger human alignment than prior automatic methods on summarization, while also warning that LLM evaluators may favor LLM-generated text.8
Those findings support automation as a scalable approximation. They do not support replacing every evaluator with the same model family that generated the behavior. Depending on task stakes and uncertainty tolerance, a robust design can vary model families, randomize response order, hide system identity, separate actor and judge, repeat stochastic runs, and measure agreement against held-out human judgments.
A practical allocation rule
The useful question is not “AI users or real users?” It is “which uncertainty deserves which kind of evidence?”
- Regression testing after every build
- Large persona and edge-case sweeps
- Prompt, policy, and conversation stress tests
- Early ranking of many design variants
- Clear, repeatable annotation rubrics
- New behavior with no calibration data
- Emotion, trust, identity, or social meaning
- Accessibility and assistive technology
- Expert or high-consequence judgments
- Claims about real adoption or causal lift
A sensible evaluation funnel is therefore:
- Broad simulation: run hundreds or thousands of controlled persona-task combinations cheaply.
- Failure-focused review: inspect disagreements, rare outcomes, and high-risk segments.
- Human calibration: compare a stratified subset with real participants and estimate where simulation fails.
- Real-world confirmation: use usability studies, field trials, or randomized experiments for claims only real behavior can support.
- Continuous monitoring: recalibrate when the product, population, model, or environment changes.
This is the workflow MatrAIx is built for: run large, persona-conditioned simulations continuously, then spend scarce human studies where they change the decision. AI simulation makes evaluation cheap enough to run before every consequential human study and after every product change. Real users then become a high-value source of calibration, discovery, and confirmation, not a scarce resource spent rediscovering obvious failures.
Reproduce the estimates
All model costs in this article follow the provider’s listed per-million-token rates accessed July 12, 2026. Human costs use public platform prices on the same date. Replace the constants below with actual measured tokens, concurrency, observed latency, traffic, and vendor terms before budgeting a project.
human = N × minutes / 60 × hourly_reward × 1.428
ai = input_tokens / 1e6 × input_rate + output_tokens / 1e6 × output_rate
hybrid = ai + calibrated_human_subset + adjudication
ai_elapsed ≈ ceil(runs / concurrency) × observed_seconds_per_run
Sources
- Prolific, Pricing; and Prolific’s payment principles. Public reward guidance and platform fees. Accessed July 12, 2026.
- User Interviews, Recruit pricing. Public per-session recruiting prices, incentive exclusions, and processing fee. Accessed July 12, 2026.
- Respondent, Participant recruitment pricing. Public session prices. Accessed July 12, 2026.
- Together AI, Serverless inference pricing. Gemma 3n E4B, Qwen3.5 9B, and DeepSeek V4 Pro token rates. Accessed July 12, 2026. Prices can change.
- Moran, “Usability (User) Testing 101,” Nielsen Norman Group; and Nielsen, “How Many Test Users in a Usability Study?” Guidance on qualitative, quantitative, and multi-group studies.
- Optimizely, “Statistical significance”. Sample size, duration, randomization, and business-cycle guidance. The numerical scenario uses a standard two-sample proportion power approximation; production tools may use different sequential methods.
- Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” NeurIPS 2023. Human agreement and documented judge biases.
- Liu et al., “G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment,” EMNLP 2023. Human correlation and evaluator-bias caveat.
- Drive Research, “How Much Does Market Research Cost?” Published ballpark budgets for online surveys, focus groups, interviews, and in-person surveys. Accessed July 12, 2026.
- Epoch AI, “LLM inference prices have fallen rapidly but unequally across tasks”; and Stanford HAI, Artificial Intelligence Index Report 2025. Inference price at a fixed capability level over time. Accessed July 12, 2026.
Method note. These are illustrative comparative estimates, not claims of measured MatrAIx accuracy, speed, or savings. They exclude taxes and shared labor; they do not price privacy, consent, environmental impact, model development, or the cost of an incorrect decision. “AI” prices cover inference only unless stated otherwise.
