Reaction
Training an autoregressive model to mimic Monte Carlo tree search
A reply calls training on node traces “super inefficient” and asks why not call a tool instead.
TLDR
One commenter asks if an autoregressive model could mimic Monte Carlo tree search by training on traces of its nodes. A reply says yes, but calls the approach “super inefficient,” comparing it to teaching a language model arithmetic when it could call a tool. The first commenter thinks chain-of-thought does something similar and agrees the efficiency is “horrible.”
Combined views
775
2 Sources, first seen ago
1 likes2 comments