Both cases appear in Figure 3(b), “Native reasoning pathways and illustrative thinking traces.” GLM-5.1 illustrates Value → Choice; Qwen3-32B illustrates Value → Uncertainty → Value → Choice. Neither case uses metacognitive prompting.
The GLM-5.1 history is documented in Appendix D, “Main-text case provenance”: Setting 2, block 3, trial 4. Arm 1 has one observed reward (7); Arm 2 has two (9, 7), giving sample means of 7 and 8. These are not the arms’ true expected rewards. The supplied Qwen3-32B excerpt does not specify its full reward history, so no history or task coordinates are reconstructed here.
The text is a selection of the paper’s excerpts, not complete raw reasoning logs.