LLM reasoning traces may lack user-interpretable meaning, even on iGSM
A researcher sharing new work says traces on iGSM, a benchmark proposed to highlight meaning in reasoning steps, lack end-user-interpretable semantics. Models trained with swapped or corrupted traces can still do well, they say.
TLDR
Earlier work made claims about the meaning of intermediate tokens on the iGSM benchmark. A researcher sharing a new paper says its traces lack end-user-interpretable semantics, and that models trained with swapped or corrupted traces can do well too.
Combined views
2.7K
1 Source, first seen 8h ago
likes