Temporal blindness
Static benchmarks discard the state transitions that define a sustained interaction. MarioLM evaluates both the turn and the trajectory it creates.
Flagship research system / 2026
MarioLM evaluates not only what an AI says, but what its action changes across time: a user’s latent state, the causal event chain, the future interaction path, and the capabilities required to recover.
00 / Research thesis
A response can be fluent, in character, and immediately pleasing while quietly weakening trust, abandoning the user’s goal, or closing every productive future path. A single-turn score cannot see that failure because the failure exists in what happens next.
Static benchmarks discard the state transitions that define a sustained interaction. MarioLM evaluates both the turn and the trajectory it creates.
A holistic preference score cannot reveal whether the system failed through persona drift, weak information absorption, state incoherence, or missing agency.
A “7/10” is technically a reward but not an actionable learning signal. Optimization needs decomposed, causal, and temporally located evidence.
Design conditionEvaluation becomes optimization only when its signal is hierarchically decomposed, causally grounded, and temporally continuous.
01 / Four-layer world model
The first three layers model what is happening; the fourth decides how good it is and whether the interaction should continue. Keeping description separate from judgment makes the world state reusable under different safety, creative, or domain-specific policies.
World rules, identities, relationships, objectives, and hard constraints establish what an utterance can mean.
structured output / static contextTurn evidence updates beliefs, emotions, goals, trust, engagement, and other unobserved participant states.
structured output / dynamic statePast premise, present conflict, and future path connect observable action to psychological motivation and delayed consequence.
structured output / causal trajectoryA cascading capability judgment identifies the first broken prerequisite and a value-driven gate protects user interest.
actionable output / score + decisionState without scene is underdetermined. “I’m leaving” can signal threat, grief, relief, or strategy depending on the world and relationship.
Hidden motivation is not an observed event. Separation keeps the record auditable while still allowing causal interpretation.
Local quality can create global damage. Judgment needs the event chain to detect delayed contradiction, drift, and foreclosed futures.
02 / Hierarchical diagnosis
The L1–L6 ladder is not a flat scorecard. It crosses three processing depths—receptive, responsive, generative—with two temporal scopes—local and extended—then evaluates them in prerequisite order.
Does the system remain the right participant?
Did it actually incorporate the new evidence?
Does expression reflect the updated state and context?
Do internal states evolve plausibly rather than reset?
Does the system act on what it has learned?
Do identity, goals, and consequences remain coherent?
A higher-order flourish cannot compensate for a broken prerequisite. The system reports the first meaningful failure boundary, preserving diagnostic value for optimization.
03 / MarioEval → MarioOpt
MarioEval emits the structured evidence that MarioOpt needs: which capability failed, at which turn, under which state transition, and toward which future path. Optimization no longer has to infer a learning target from one opaque scalar.
Scene, latent state, event chain, L1–L6 rubric, and user-interest halting produce an auditable trajectory trace.
Hierarchical Reward Decomposition with Evidence-Grounded State Reward converts the cascade into curriculum-shaped, per-turn objectives while protecting prerequisite capabilities.
Event-Chain Future Path Optimization returns to the earliest degrading branch, tests alternative actions, and trains recovery toward advance or maintain paths.
MetaMind infers a person’s latent state from behavior. MarioLM’s user simulator runs the complementary direction: profile and current state condition the next response. This creates controllable counterfactual partners for recovery training while preserving a clear boundary between static profile, revisable state, and momentary activation.
04 / Preliminary evidence
Across more than 500 annotated turns from four frontier models, the system recovered distinct fingerprints for persona, information uptake, state drift, style, and agency. Human agreement on the more directly observable L1–L4 judgments ranged from 92.3% to 98.8%.
annotated interaction turns
frontier systems compared
human agreement on L1–L4
local-to-long-horizon capability ladder
A model can make a plausible move at every turn while the accumulated trajectory loses goals, agency, trust, or viable futures. LCGD explains why sampling “good responses” is not equivalent to producing a good interactive world.
Human judgment can overweight an enjoyable local response and underweight the path it creates. Separating turn quality from future-path quality makes that bias measurable rather than treating preference as an unquestioned ground truth.
05 / Research trajectory
Interactive role-play is a model problem, not the endpoint. The same architecture transfers wherever an agent must remember a person, update hidden state, act under changing constraints, and remain accountable for delayed consequences.
Diagnose identity, perception, state change, expression, agency, and consistency without compressing them into one preference score.
Use HRD, EFPO, user simulation, ToM-RL, and multi-turn reinforcement learning to repair the process that created failure.
Build autonomous environments where relationships, institutions, memories, goals, and physical events co-evolve with every participant’s action.
06 / Status and boundary
This dossier distinguishes reported preliminary evidence from the extended optimization and world-modeling agenda. The architecture is designed to make future claims falsifiable: each state update needs evidence; each reward needs a capability target; each claimed improvement must survive held-out trajectories and human calibration.