MatrAIx builds structured user representations that power simulation and evaluation infrastructure.
Our internal corpus combines synthetic generation, evidence-aware extraction
from human-authored public data, and consented self-reports, all mapped into one inspectable schema.
The corpus is not a single undifferentiated dataset. Generation method, source context, confidence, and evidence boundaries matter.
01 · SYNTHETIC
Structured generation
Personas are sampled from a typed attribute graph, with compatibility rules designed to reduce contradictory combinations and preserve diversity.
Schema-constrained attributes
Conditional consistency rules
Reproducible sampling and rendering
02 · HUMAN-EVIDENCE
Grounded human data
Public traces and consented survey responses are mapped to supported schema attributes. Unobserved fields should remain unknown rather than being invented.
Evidence or self-report attached to assignments
Confidence, consent, and provenance tracking
Source-specific collection policies
Human persona sources
Different sources, different signals.
01
Wikipedia
Public biographical and professional context, mapped only where the source provides support.
BIOGRAPHICAL02
Amazon Reviews
Preference, product-use, and decision signals inferred from review histories with explicit evidence boundaries.
BEHAVIORAL03
Stack Overflow
Technical interests, tools, expertise, and developer-workflow signals derived from public contributions.
TECHNICAL04
Persona Survey
Consented, first-person responses provided directly by participants and mapped to the shared persona schema.
SELF-REPORTED
No source makes every attribute observable. Sensitive or unsupported claims are excluded, and collection or release eligibility depends on consent, source terms, privacy review, and documentation.
8.3B
Release policy
Research at scale. Release with restraint.
We will not directly distribute the full 8.3B internal corpus. It includes large-scale synthetic outputs, human-evidence extraction artifacts, and consented self-reports that require source-aware governance, storage, and quality controls.
PLANNED PUBLIC DATASET1,000,000 personas
A curated subset is planned for Hugging Face with a dataset card, schema documentation, provenance summaries, validation notes, and stated limitations.