Home/Research/Publication

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

A behavioral-science instrument for measuring how persona-conditioned agents revise their identities under sustained environmental pressure.

Accepted COLM 2026 @ SocialSim
MatrAIx Research Community
July 26, 2026
4 min read
Twenty-five agents moving through the resource-scarce MicroVerse desert environment
Twenty-five agents inhabit a deterministic desert world where water scarcity creates sustained survival pressure.

Long-horizon social simulations often assume that persona-conditioned agents remain faithful to their identities. MicroVerse turns that assumption into something we can measure.

Accepted to the COLM 2026 Workshop on Social Simulation with LLMs, MicroVerse is a behavioral-science instrument for studying self-authored identity drift. Each agent begins with an immutable original “soul file” containing values, moral boundaries, personality, and goals, alongside a mutable current identity that can change only through the agent's own reflection.

Pressure with measurable consequences

Twenty-five agents inhabit a deterministic 50 × 50 desert environment. Water drains every tick, the central Atmospheric Siphon cannot supply everyone, and death is permanent. The eight available actions, including trade, talk, attack, and scavenging, directly engage the agents' moral boundaries.

The instrument separates revision from measurement. Agents decide organically whether reflection has genuinely changed them. The engine independently records identity snapshots at fixed intervals and captures every agent at the end, including agents that died, reducing sampling and survivor bias.

MicroVerse measures the structural gap between an immutable starting identity and an agent's self-authored current identity, rather than relying on a single opaque similarity score.

Two preliminary findings

First, anti-self-deception emerged without being prompted. Across the pilot, 27 of 111 added moral boundaries described agents recognizing and rejecting their own rationalizations, the largest semantic category of identity modification.

Second, a reflection-threshold sweep at 40, 80, and 150 showed that the gate controls when and how often revisions occur, but not their qualitative direction. Lower thresholds produced earlier and more frequent revisions; where ruthless personas drifted, they consistently acquired guardrails toward others.

These findings are preliminary existence proofs from one model and limited seeds, not population-level significance claims. The next experiments will add no-persona baselines, memory and framing ablations, stronger censoring adjustments, multi-judge reliability, and broader model and seed coverage.