Evolution Strategies for LLM Post-Training Examined
A retweet shares research on using Evolution Strategies to refine large language models without gradients.
Yee Whye Teh retweeted a post by @zhengzhi20 that presents a study on Evolution Strategies for LLM reasoning. The post states that ES enables post-training of large language models without backpropagation. It claims the work finds Evolution Strategies achieve higher Pass@K than GRPO during LLM post-training. The accompanying summary describes ES as a gradient-free method that maintains broader exploration compared with other approaches. The discussion centers on this alternative training technique and its reported results for model refinement.
Combined views
14
1 post, first seen 3h ago
Evolution Strategies for LLM Post-Training Examined
A retweet shares research on using Evolution Strategies to refine large language models without gradients.