Does an AI’s chain of thought faithfully reflect how it reached an answer?
A thread examines two papers on what written-out AI reasoning reveals: which steps matter to an answer, and whether a model keeps generating reasoning after committing to one.
TLDR
A user questions whether chain-of-thought text faithfully records how an AI produces an answer, citing two papers. The thread describes “Legibility is Not Interpretability” as testing whether AI judges can identify which reasoning steps mattered. The judges reportedly beat chance but remained far below the achievable ceiling, especially for correct solutions. The thread also describes “Drop the Act,” which studies “reasoning theater”—reasoning generated after a model has internally committed to an answer. According to the thread, using a signal from the model’s internal activations during reinforcement learning reduced measured reasoning theater by 11–100%, depending on the setting. Chains of thought became up to 19% shorter, while accuracy stayed roughly unchanged.
Combined views
1.1K
5 Sources, first seen 1d ago
Does an AI’s chain of thought faithfully reflect how it reached an answer?
A thread examines two papers on what written-out AI reasoning reveals: which steps matter to an answer, and whether a model keeps generating reasoning after committing to one.
TLDR
A user questions whether chain-of-thought text faithfully records how an AI produces an answer, citing two papers. The thread describes “Legibility is Not Interpretability” as testing whether AI judges can identify which reasoning steps mattered. The judges reportedly beat chance but remained far below the achievable ceiling, especially for correct solutions. The thread also describes “Drop the Act,” which studies “reasoning theater”—reasoning generated after a model has internally committed to an answer. According to the thread, using a signal from the model’s internal activations during reinforcement learning reduced measured reasoning theater by 11–100%, depending on the setting. Chains of thought became up to 19% shorter, while accuracy stayed roughly unchanged.