Research Scientist · Qwen

Intelligence is what survives
after the next token.

I build AI systems that expose hidden computation, reason about other minds, and remain coherent when actions accumulate into consequences.

Xuanming Zhang
Research frame 01Human-centered frontier models
941
Scholar citations
Spotlight
NeurIPS 2025
ACL · ICLR
2026
Qwen · Stanford · Amazon
Research trajectory

Research premise

Most benchmarks ask whether a model can answer.
I ask what it computes, what it understands, and what its answer changes.

One program

Mechanism discovery → cognitive structure → experiential evaluation → optimization.

01 / Research instrument

Three questions.
One research program.

Use the instrument to inspect the same intelligence problem at three scales: computation inside the model, cognition between agents, and consequences across time.

Confident Decoding

The final layer is not always the most truthful layer.

Aligned models repeatedly move through Guess → Refine → Perturb. Scrub the model depth to see where reasoning peaks before the final-layer alignment tax.

Read the mechanism paper ↗
InputModel depthOutput
123456789101112131415161718
GuessRefinePerturb
Layer 14 · reasoning refinement

02 / Selected systems

Research that leaves
a working machine behind.

Case 01Inference systems · Qwen

Confident Decoding

Recover the reasoning that alignment perturbs.

A training-free, entropy-guided layer-selection method grounded as an optimal-stopping problem and integrated into a production vLLM stack.

  • +6.5 GPQA-D points
  • 0 MB KV-cache overhead
  • <2% latency
Case 02Cognitive systems · NeurIPS Spotlight

MetaMind → CooT → Latent-RL

Move cognition from scaffold to evolvable capability.

A continuous research lineage: make mental states explicit, make alignment auditable at decoding time, then internalize the process through reinforcement learning.

  • +35.7% social scenarios
  • 16+ models evaluated
  • 0.80 AIR-Bench
Case 03World modeling · Research grant

MarioLM

Turn evaluation into an experiential world.

MarioLM unifies trajectory-level evaluation and optimization through causally grounded scenes, latent user states, event chains, and temporally persistent consequences.

  • 500+ annotated turns
  • 92.3–98.8% human agreement
  • L1–L6 cascaded rubrics
Trajectory audit / illustrative traceThe response still looks good.
The relationship does not.

Single-turn quality remains above 90 while trust and agency deteriorate across repeated interaction.

Local response quality compared with relational trust and user agency over six turnsLocal response quality changes from 94 to 90, relational trust from 92 to 34, and user agency from 91 to 23.1007550250T1T2T3T4T5T6trajectory warning
Local response quality94 → 90Relational trust92 → 34User agency91 → 23
Turn-level judge6 / 6 pass
Trajectory judgeFailure visible at T4

The object of evaluation is not an isolated answer, but the state transition induced by a sequence of answers.

LH-Deception · ICLR 2026

Audit the chain, not the confession.

Controlled performer–supervisor–auditor worlds reveal concealment and fabrication that do not exist in single-turn tests.

Paper ↗

Trustworthy code intelligence

Keep evaluation fresh, realistic, and difficult to game.

DevEval, EvoCodeBench, CDD/TED, and Seeker connect repository realism, benchmark freshness, black-box contamination control, and exception-safe generation.

03 / Selected publications

Evidence, not inventory.

EMNLP 2026Under review · top 1%

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

Xuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang, et al.

NeurIPS 2025Spotlight

MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

ACL 2026Long paper

Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

ICLR 2026Equal first author

Simulating and Understanding Deceptive Behaviors in Long-Horizon Interactions

Yang Xu*, Xuanming Zhang*, Samuel Yeh, Jwala Dhamala, Ousmane Dia, Rahul Gupta, Sharon Li

PNAS 2026Human subjects

Human Preferences Are Susceptible to Covertly Misaligned AI Advice

Sahand Sabour, June M. Liu, Siyang Liu, Chris Z. Yao, Shiyao Cui, Xuanming Zhang, et al.

Complete publication record on Google Scholar ↗
A computational research atlas connecting model mechanisms, social cognition, and long-horizon worlds

04 / Field notes

The paper reports the result.
The field note reveals the thinking.

An evidence-led technical account and a clearly marked research agenda trace MetaMind from metacognitive social reasoning toward Cognitive AI.

Open the MetaMind dossier →

05 / Trajectory

From reliable code
to persistent minds.

  1. Qwen, Alibaba GroupResearch Scientist · A Star Top Talent

    Subjective capabilities, test-time reasoning systems, multi-turn interaction, and mental-state reinforcement learning.

  2. Stanford NLP GroupVisiting Student Researcher

    Human-centered AI and social reasoning under uncertainty.

  3. Amazon AGIStudent Researcher

    Long-horizon deception, trust dynamics, and foundation-model safety evaluation.

  4. PKU · UW–Madison · ByteDance · Zhipu AIResearch and systems

    Trustworthy code evaluation, multimodal agents, psychological simulation, and metacognitive social intelligence.