Reaction
Chain-of-thought framed as learned search on an append-only tape
A post says the same model generates and evaluates, without an explicit value function at inference.
TLDR
A post frames chain-of-thought (CoT) as learned search serialized onto an append-only tape. It says the search is amortized during training and that using one sequence makes it mostly depth-first. When the model backtracks, the post notes, dead branches stay in context.
Combined views
296
1 Source, first seen ago
8 likes1 reposts