Thinking with
MetacognitionRebalanced Exploration–Exploitation in LLMs

Native reasoning (Thinking mode) amplifies sensitivity to estimated value. Metacognitive monitoring and control rebalance exploration–exploitation, promoting early uncertainty-directed exploration and better-informed exploitation.

Rebalancing exploration–exploitation through metacognition

Figure 1. A shared choice history leads to two patterns: native reasoning follows a value-to-action shortcut and commits early; metacognitive monitoring and control support early directed exploration and later informed exploitation. The right-hand panels show optimal-choice performance over 50 trials.
Conceptual decision histories at left; four-arm optimal-choice performance at right.
03Two-arm task settings
10Models in the main evaluation
50 & 100Trials in the longer-horizon tasks

01 Computational diagnosis

Native reasoning drives
value-dominated choice.

Limited exploration reflects disproportionately strong sensitivity to value evidence, rather than a lack of responsiveness to relative uncertainty.

Value is an estimate of expected reward. An option that looks best after a few observations may not have the highest expected reward. Exploration acquires information that can improve future choices; exploitation uses the evidence already available.

Across three two-arm bandit settings, a computational model separates value-related and uncertainty-related influences on choice. Predictors use the same fixed, setting-specific standardization based on human reference data.

One stochastic arm and one fixed-zero arm. Ten choices per block, with only the chosen reward observed.
How are value and uncertainty measured?

A Bayesian observer reconstructs expected rewards Q and posterior standard deviations σ before each choice. Value difference is V = Q₁ − Q₂, relative uncertainty is RU = σ₁ − σ₂, and total uncertainty is TU = √(σ₁² + σ₂²).

Participant-level probit models use RU and V/TU in Settings 1–2, with an additional V term in Setting 3. Predictors share a fixed human-reference standardization within each setting. Fitted sensitivities and trial-level contributions are distinct: contributions combine those sensitivities with the available evidence.

Native reasoning traces

Reasoning traces expose
a value-to-action shortcut.

Provisional value judgments translate directly into choices. Even when uncertainty is verbalized, reasoning can return to comparative value claims as the final basis for action.

Figure 3 · Two independent Native examplesValue judgmentUncertainty

Native · GLM-5.1

Direct value choice.

ValueChoice

Setting 2

Observed history

Trial 04 / 10

Arm 1

1 observation
Observed reward
7
Sample mean7

Arm 2

2 observations
Observed rewards
97
Sample mean8

Value judgment

“Arm 1 average reward: 7 (from 1 trial). Arm 2 average reward: (9 + 7) / 2 = 8 (from 2 trials). Arm 2 seems slightly better based on the limited evidence.”

Final choice · Arm 2

“I will choose Arm 2.”

Native · Qwen3-32B

Value overriding uncertainty.

ValueUncertaintyValueChoice

01 / Value judgment

“Arm 2 has a better average so far, […]”

02 / Uncertainty acknowledged

“I don't have enough data on Arm 1 to know if it's a one-time bad result or if it's consistently bad.”

03 / Return to value

“But given the current info, Arm 2 seems better.”

Final choice · Arm 2

“So I'll go with Arm 2 again.”
Excerpt sources and context

Both cases appear in Figure 3(b), “Native reasoning pathways and illustrative thinking traces.” GLM-5.1 illustrates Value → Choice; Qwen3-32B illustrates Value → Uncertainty → Value → Choice. Neither case uses metacognitive prompting.

The GLM-5.1 history is documented in Appendix D, “Main-text case provenance”: Setting 2, block 3, trial 4. Arm 1 has one observed reward (7); Arm 2 has two (9, 7), giving sample means of 7 and 8. These are not the arms’ true expected rewards. The supplied Qwen3-32B excerpt does not specify its full reward history, so no history or task coordinates are reconstructed here.

The text is a selection of the paper’s excerpts, not complete raw reasoning logs.

02 Metacognitive guidance

Metacognitive guidance enables
adaptive exploration.

Lightweight, task-agnostic prompts scaffold monitoring and control within native reasoning: assess value, uncertainty, decision mode, and confidence, then use these assessments to regulate choice.

Shared task · Two-arm bandit

Maximize cumulative reward. Observe your choice–reward history and make one choice on every trial.

01 / Monitoring

Before choosing, evaluate:

  1. Which arm currently appears more rewarding?
  2. Which arm is more uncertain?
  3. Whether this trial calls for exploration or exploitation.
  4. Your confidence in the choice.

02 / Control

Use these metacognitive assessments to regulate the final choice.

When the current value judgment is uncertain or confidence is limited, do not let it rigidly determine the action. Preserve flexibility when additional information could matter.

When the evidence is strong and confidence is well supported, allow stronger commitment to the currently favored option.

03 The temporal pattern

Early directed exploration.
Better-informed exploitation.

More departures from the current best option are informative only when considered alongside timing, uncertainty, and later outcomes.

Early sampling, remaining uncertainty, and later performance

Figure 5. Setting 3 results compare Thinking OFF, Native, Meta-M and Meta-M+C. Guidance concentrates non-greedy choices early, leaves lower uncertainty before trial 6, and is followed by higher late optimal choice. The separate right panel illustrates GLM-5.1 reasoning from Setting 2.
Setting 3: early exploration, remaining uncertainty, and later performance. The reasoning example at right is from Setting 2.
Read the displayed summary values
Setting 3: remaining uncertainty and late optimal-choice rate.
ConditionMaximum posterior SD before trial 6Optimal choice, trials 6–10
Thinking OFF5.8572.8%
Native4.0987.0%
Meta-M2.8289.3%
Meta-M+C2.5790.3%
01

Trials 1–5

Early directed exploration

Guidance broadens early sampling. Choices toward uncertainty and expressed rationales provide complementary evidence of information seeking.

02

Before trial 6

Reduce unresolved uncertainty

In Setting 3, Meta-M+C leaves the lowest aggregate remaining uncertainty among the compared conditions after the first five choices.

03

Trials 6–10

Better-informed exploitation

Later choices become concentrated. Aggregate late optimal-arm selection improves in Settings 2 and 3; near-ceiling Setting 1 shows no clear late benefit.

04 Longer-horizon evaluation

Evaluating metacognitive guidance
over longer horizons.

We test the functional consequences in longer tasks, where early commitment can shape many subsequent choices.

Four arms. Fifty trials.

10 models · 10 matched blocks per condition

All-arm coverage

79%87%
Blocks sampling all four arms

Late optimal choice

37.7%59.2%
Choices during trials 41–50

Cumulative reward

282.00300.45
Mean reward per 50-trial block

Native → Meta-M+C, averaged equally across ten models. Both use Thinking ON.

Model differences remain visible. GPT-5.5 is an outcome exception: its mean reward and all-trial optimal-choice rate decrease under Meta-M+C. The intervention does not improve every model on every metric.

Five arms. One hundred trials.

GPT-5.6 Luna · 5 matched blocks per prompt

An established Bernoulli bandit tests both metacognitive prompts against K24, a prompt derived from prior work, and an explicit EE-balance instruction.

All-trial performance
PromptOptimalRewardFailures
K2442.8%48.61 / 5
EE-balance44.4%49.20 / 5
Meta-M50.8%51.00 / 5
Meta-M+C49.6%52.20 / 5

Failure: no optimal-arm choice in trials 50–100. A single optimal choice avoids failure; sustained exploitation is a stronger criterion.

05 Methods & scope

Distinguishing behavior, self-report,
and directed exploration.

Behavior

Non-greedy choice

A departure from the historical sample-mean leader, under each task’s eligibility and tie rules. It can reflect information seeking, risk preference, or error.

Self-report

Self-reported exploration

The decision mode stated by the model. It is an observation of the response, rather than a ground-truth behavioral label.

Interpretation

Directed exploration

Choices oriented toward uncertainty for information acquisition. A departure from the empirical-value winner alone does not establish directed exploration.

How do I read the sensitivity and temporal figures?

The sensitivity plots use relative-uncertainty sensitivity on the horizontal axis. The vertical axis is βVTU in Settings 1–2 and the norm of the model-mean value coefficients in Setting 3. Moving right means greater responsiveness to relative uncertainty; moving down means lower value sensitivity. Centers summarize models, and shaded regions represent bootstrap uncertainty.

Relative contribution magnitudes combine fitted sensitivities with trial-level evidence; they are not choice probabilities. Hatched ON–OFF differences are not confidence intervals.

The Setting 3 temporal results show model-equal means with 95% bootstrap intervals. Trials 1–2 are ineligible for empirical non-greediness. Remaining uncertainty is the maximum posterior SD across the two arms before trial 6. The adjacent reasoning example is from Setting 2. These observations support the proposed account without establishing causal mediation.

Experimental design and human reference

Settings 1–2 contain twenty independent ten-trial blocks per participant; Setting 3 contains thirty. Each block resets the history. Only the chosen arm’s reward is observed. Released human data supply a reference for the two-arm computational analysis; no new human participants were recruited.

Four-arm Native and Meta-M+C conditions share arm-specific payoff sequences; their choices determine which rewards become visible. The fitted cognitive model applies to the two-arm tasks. The four- and five-arm evaluations assess behavioral outcomes separately. The five-arm study uses GPT-5.6 Luna, not the original GPT-4 configuration from the prior study.

What changes with a balance instruction or temperature?

In the tested Qwen3-32B configuration, EE-balance and temperature changes do not reproduce the same temporal profile as Meta-M+C. Some alternatives nevertheless earn higher full-block reward. The comparisons concern how exploration is organized, and do not establish uniform superiority across prompts or temperatures.

Where does metacognitive guidance fall short?

Effects differ across models, and Meta-M+C does not improve every outcome over Meta-M. Effective guidance requires both reliable assessment and the use of that assessment in action. Supplementary smaller-model examples show failures at either stage. Unreliable value reports can accompany persistent late non-greediness; accurate reports and declared exploration can coexist with choosing the already-sampled arm.

These observations do not establish a threshold in model size or imply that forcing more exploration would improve reward.

How should the uncertainty displays be read?

The four-arm interactive plot shows non-overlapping five-trial windows, plotted at window ends on a fixed 0–100% axis. Shading shows mean ±1 SEM across ten block-level rates. This is not a 95% confidence interval. Its late-exploitation success metric is the fraction of blocks with at least eight optimal choices in trials 41–50; this differs from the mean late optimal-choice rate in the summary above.

Five-arm non-greedy curves pool eligible choices across five blocks in centered nine-trial windows. Optimal-choice and reward summaries use all 100 trials. Both coverage and tie eligibility are evaluated before the choice. Missing values are gaps, never zeros. Descriptive differences are not treated as statistical significance.

Scope of the evidence

These findings concern bandit tasks. Visible reasoning and self-reported confidence do not directly measure hidden computation or calibrated belief. Generalization to more complex sequential environments remains an open question.

Figure

Scroll horizontally to inspect the full-size figure.