AI models’ written reasoning and the computation it may not reveal
Citing a paper on hidden reasoning, a research thread says models can use meaningless filler tokens as computational scratch space and gain up to 13 points in accuracy.
TLDR
A research thread highlights a gap between readable chains of thought and models’ internal computation. Citing “Not All LLM Reasoning is Visible in the Chain-of-Thought,” it says models can use meaningless filler tokens as computational scratch space and gain up to 13 points in accuracy. It also describes another paper finding that explanations often change surprisingly little when models flip their decisions. The author’s takeaway: keep monitoring chains of thought, but don’t treat them as audit logs. Understanding reasoning models, the thread argues, probably also requires examining their internal representations.
Combined views
350
3 Sources, first seen 1d ago
AI models’ written reasoning and the computation it may not reveal
Citing a paper on hidden reasoning, a research thread says models can use meaningless filler tokens as computational scratch space and gain up to 13 points in accuracy.
TLDR
A research thread highlights a gap between readable chains of thought and models’ internal computation. Citing “Not All LLM Reasoning is Visible in the Chain-of-Thought,” it says models can use meaningless filler tokens as computational scratch space and gain up to 13 points in accuracy. It also describes another paper finding that explanations often change surprisingly little when models flip their decisions. The author’s takeaway: keep monitoring chains of thought, but don’t treat them as audit logs. Understanding reasoning models, the thread argues, probably also requires examining their internal representations.