Zhi Zheng Posts Evolution Strategies Study for LLMs
Ph.D. candidate at National University of Singapore posts study summary on X with attached research figure.
TLDR
Zhi Zheng, a Ph.D. candidate at National University of Singapore, posted on X about a study on Evolution Strategies for LLM reasoning. The post states that ES can post-train LLMs without backprop. It reports that ES explores higher Pass@K than GRPO, produces sparse functional updates with no catastrophic forgetting, and requires fewer samples as LLMs grow larger. The tweet includes a single research figure labeled as Figure 1 with three panels. The message cuts off while noting that the paper is summarized through three research points.
Combined views
14.4K
1 Source, first seen 31d ago