Board-game self-play method is claimed to transfer LLM reasoning to unseen games and math
A post describes Self-Play Search Distillation as turning the search of MuZero experts trained through board-game self-play into reasoning traces for LLMs.
TLDR
The post claims those reasoning traces transfer to unseen games and math. It gives a math figure that rises from 24.1 to 36.6, without specifying what the figure measures.
Combined views
816
2 Sources, first seen 4h ago
31 likes
