Use the instrument to inspect the same intelligence problem at three scales: computation inside the model, cognition between agents, and consequences across time.
Confident Decoding
The final layer is not always the most truthful layer.
Aligned models repeatedly move through Guess → Refine → Perturb. Scrub the model depth to see where reasoning peaks before the final-layer alignment tax.
Select a competing mental-state hypothesis. The system makes latent interpretations explicit, checks them against social priors, then uses the surviving state to guide its response.
Polite reluctanceSurface consent conflicts with contextual hesitation.
MarioLM · LH-Deception
A good next token can still create a bad future.
Move through a five-step interaction. Local quality remains plausible while latent intent, trust, and agency drift—failure modes that only become visible at trajectory scale.
Move cognition from scaffold to evolvable capability.
A continuous research lineage: make mental states explicit, make alignment auditable at decoding time, then internalize the process through reinforcement learning.
MarioLM unifies trajectory-level evaluation and optimization through causally grounded scenes, latent user states, event chains, and temporally persistent consequences.