MatrAIxGitHub ↗Join us
// PERSONA DATA & SCHEMA loading ACTIVE RESEARCH

Personas at planetary scale.

MatrAIx builds structured user representations that power simulation and evaluation infrastructure. Our internal corpus combines synthetic generation, evidence-aware extraction from human-authored public data, and consented self-reports, all mapped into one inspectable schema.

8.3Binternal persona corpus
1Mplanned public release
schema dimensions
schema categories
How we build

Two pipelines. One evidence-aware schema.

The corpus is not a single undifferentiated dataset. Generation method, source context, confidence, and evidence boundaries matter.

01 · SYNTHETIC

Structured generation

Personas are sampled from a typed attribute graph, with compatibility rules designed to reduce contradictory combinations and preserve diversity.

  • Schema-constrained attributes
  • Conditional consistency rules
  • Reproducible sampling and rendering
02 · HUMAN-EVIDENCE

Grounded human data

Public traces and consented survey responses are mapped to supported schema attributes. Unobserved fields should remain unknown rather than being invented.

  • Evidence or self-report attached to assignments
  • Confidence, consent, and provenance tracking
  • Source-specific collection policies
Human persona sources

Different sources, different signals.

01

Wikipedia

Public biographical and professional context, mapped only where the source provides support.

BIOGRAPHICAL
02

Amazon Reviews

Preference, product-use, and decision signals inferred from review histories with explicit evidence boundaries.

BEHAVIORAL
03

Stack Overflow

Technical interests, tools, expertise, and developer-workflow signals derived from public contributions.

TECHNICAL
04

Persona Survey

Consented, first-person responses provided directly by participants and mapped to the shared persona schema.

SELF-REPORTED

No source makes every attribute observable. Sensitive or unsupported claims are excluded, and collection or release eligibility depends on consent, source terms, privacy review, and documentation.

Release policy

Research at scale. Release with restraint.

We will not directly distribute the full 8.3B internal corpus. It includes large-scale synthetic outputs, human-evidence extraction artifacts, and consented self-reports that require source-aware governance, storage, and quality controls.

PLANNED PUBLIC DATASET1,000,000 personas

A curated subset is planned for Hugging Face with a dataset card, schema documentation, provenance summaries, validation notes, and stated limitations.

Hugging Face · coming soon Follow development ↗
Release timing, license, composition, and final counts are placeholders pending review.
Live schema explorer

Inspect every dimension.

dimensions across categories, with allowed values. Search, filter, expand, or sample a structured persona below.

Canonical schema ↗