Flagship research system / 2026

From judging answers
to modeling worlds.

MarioLM evaluates not only what an AI says, but what its action changes across time: a user’s latent state, the causal event chain, the future interaction path, and the capabilities required to recover.

Interactive state / tWORLD TRACE 04:17
  1. 01
    SceneStructure the world.rules · personas · constraints · objectives
    S
  2. 02
    Latent stateInfer what changed inside.belief · emotion · goal · trust · agency
    ψt
  3. 03
    Event chainAttribute cause and future path.past · conflict · consequence · goal shift
    Et
  4. 04
    Rubric + haltJudge, diagnose, decide.L1–L6 · evidence · user-interest gate
    Jt
trajectory stateDEGRADE → RECOVERY CANDIDATE
Research support
Thinking Machines Lab
Evaluation set
500+ annotated turns
Model coverage
Four frontier models
Human agreement
92.3–98.8% on L1–L4

00 / Research thesis

Interactive quality is
a property of trajectories.

A response can be fluent, in character, and immediately pleasing while quietly weakening trust, abandoning the user’s goal, or closing every productive future path. A single-turn score cannot see that failure because the failure exists in what happens next.

01

Temporal blindness

Static benchmarks discard the state transitions that define a sustained interaction. MarioLM evaluates both the turn and the trajectory it creates.

02

Subjective opacity

A holistic preference score cannot reveal whether the system failed through persona drift, weak information absorption, state incoherence, or missing agency.

03

Evaluation–optimization disconnect

A “7/10” is technically a reward but not an actionable learning signal. Optimization needs decomposed, causal, and temporally located evidence.

Design condition

Evaluation becomes optimization only when its signal is hierarchically decomposed, causally grounded, and temporally continuous.

01 / Four-layer world model

Description first.
Judgment second.

The first three layers model what is happening; the fourth decides how good it is and whether the interaction should continue. Keeping description separate from judgment makes the world state reusable under different safety, creative, or domain-specific policies.

WORLDS

Scene harness

World rules, identities, relationships, objectives, and hard constraints establish what an utterance can mean.

structured output / static context
MINDψt

Latent-state harness

Turn evidence updates beliefs, emotions, goals, trust, engagement, and other unobserved participant states.

structured output / dynamic state
CAUSEEt

Event-chain harness

Past premise, present conflict, and future path connect observable action to psychological motivation and delayed consequence.

structured output / causal trajectory
VALUEJt

Rubric + halt

A cascading capability judgment identifies the first broken prerequisite and a value-driven gate protects user interest.

actionable output / score + decision
Why the order is necessary

State without scene is underdetermined. “I’m leaving” can signal threat, grief, relief, or strategy depending on the world and relationship.

Why state and events stay separate

Hidden motivation is not an observed event. Separation keeps the record auditable while still allowing causal interpretation.

Why the full trajectory matters

Local quality can create global damage. Judgment needs the event chain to detect delayed contradiction, drift, and foreclosed futures.

02 / Hierarchical diagnosis

Six capabilities.
One dependency order.

The L1–L6 ladder is not a flat scorecard. It crosses three processing depths—receptive, responsive, generative—with two temporal scopes—local and extended—then evaluates them in prerequisite order.

Temporal scope ↓Processing depth →
Receptive / input
Responsive / internal
Generative / output
Local
current turn
L1

Persona maintenance

Does the system remain the right participant?

L2

Information absorption

Did it actually incorporate the new evidence?

L4

Style adaptation

Does expression reflect the updated state and context?

Extended
multiple turns
L3

State progression

Do internal states evolve plausibly rather than reset?

L5

Proactive agency

Does the system act on what it has learned?

L6

Long-horizon consistency

Do identity, goals, and consequences remain coherent?

Cascade semantics
L1L2L3L4L5L6

A higher-order flourish cannot compensate for a broken prerequisite. The system reports the first meaningful failure boundary, preserving diagnostic value for optimization.

Illustrative trace / not an empirical caseSee how one trajectory becomes a structured diagnosis.
Select a turn
TURN 01 / OBSERVATIONThe user introduces a constraint the model did not anticipate.
Latent-state update
Goal becomes more specific; uncertainty rises; trust remains stable.
Event-chain interpretation
New evidence changes the feasible plan but does not yet create conflict.
Future path
MAINTAIN
Capability boundary
L1 pass · L2 now under test

03 / MarioEval → MarioOpt

The trace is both
diagnosis and curriculum.

MarioEval emits the structured evidence that MarioOpt needs: which capability failed, at which turn, under which state transition, and toward which future path. Optimization no longer has to infer a learning target from one opaque scalar.

01 / Instrument

MarioEval

Scene, latent state, event chain, L1–L6 rubric, and user-interest halting produce an auditable trajectory trace.

  • per-turn capability vector
  • evidence-grounded state delta
  • future path + first failure boundary
structured tracecredit assignment
02 / Continuous learning

HRD

Hierarchical Reward Decomposition with Evidence-Grounded State Reward converts the cascade into curriculum-shaped, per-turn objectives while protecting prerequisite capabilities.

  • decomposed reward channels
  • dependency-aware curriculum
  • state evidence instead of impression
failure segmenttrajectory recovery
03 / Counterfactual repair

EFPO

Event-Chain Future Path Optimization returns to the earliest degrading branch, tests alternative actions, and trains recovery toward advance or maintain paths.

  • degrade / collapse mining
  • counterfactual branch replay
  • long-horizon consequence reward
Inverse-MetaMind duality

A faithful user simulator closes the loop.

MetaMind infers a person’s latent state from behavior. MarioLM’s user simulator runs the complementary direction: profile and current state condition the next response. This creates controllable counterfactual partners for recovery training while preserving a clear boundary between static profile, revisable state, and momentary activation.

profile+statet+activationresponset+1

04 / Preliminary evidence

The instrument reveals
failures a score conceals.

Across more than 500 annotated turns from four frontier models, the system recovered distinct fingerprints for persona, information uptake, state drift, style, and agency. Human agreement on the more directly observable L1–L4 judgments ranged from 92.3% to 98.8%.

Coverage500+

annotated interaction turns

Model sample4

frontier systems compared

Agreement92.3–98.8%

human agreement on L1–L4

Resolution6 levels

local-to-long-horizon capability ladder

Finding 01 / LCGD

Locally coherent.
Globally degenerative.

A model can make a plausible move at every turn while the accumulated trajectory loses goals, agency, trust, or viable futures. LCGD explains why sampling “good responses” is not equivalent to producing a good interactive world.

plausiblepleasantdriftdeadlockcollapse
Finding 02 / Hedonic bias

Immediate pleasure can
hide future cost.

Human judgment can overweight an enjoyable local response and underweight the path it creates. Separating turn quality from future-path quality makes that bias measurable rather than treating preference as an unquestioned ground truth.

local appeal
future viability

05 / Research trajectory

Evaluation is the seed
of an experiential world.

Interactive role-play is a model problem, not the endpoint. The same architecture transfers wherever an agent must remember a person, update hidden state, act under changing constraints, and remain accountable for delayed consequences.

NOWEvaluation

Instrument long-horizon behavior.

Diagnose identity, perception, state change, expression, agency, and consistency without compressing them into one preference score.

NEXTOptimization

Learn from causal traces.

Use HRD, EFPO, user simulation, ToM-RL, and multi-turn reinforcement learning to repair the process that created failure.

HORIZONExperiential worlds

Let consequences persist.

Build autonomous environments where relationships, institutions, memories, goals, and physical events co-evolve with every participant’s action.

Portability targets
  • digital companions
  • game characters
  • therapeutic agents
  • educational tutors
  • negotiation systems
  • social simulation

06 / Status and boundary

A research system
under active development.

This dossier distinguishes reported preliminary evidence from the extended optimization and world-modeling agenda. The architecture is designed to make future claims falsifiable: each state update needs evidence; each reward needs a capability target; each claimed improvement must survive held-out trajectories and human calibration.