MatrAIx Eval InfraPreview
Evaluation task

Simulation live Eval Net API --:--:--
Agent Swarm— active
■ shows the current batch ratio
▸ In focus · 8 agents · scroll
Evaluation metrics● LIVE
● pass ● flag ● fail
Metrics900+ or your own
Each agent run is scored against selected metrics
Aggregated result · pass / flag / fail
active agents
behaviors / sec
mean score
pass rate
Reports
Trajectory Telemetryobs → act → reward
— GENERATED REPORTS

Reports from eight evaluation tasks.

Public demo results with independent randomized sampling for every task.

2 Surveys · 1M responses 2 Chatbots · 5.2B tokens 2 Websites · 3.6M interactions 2 Apps · 54.2K actions