Specific text cues reportedly boost base models on MATH-500
A research post says opening tokens sharply raised Olmo-3-7B and Qwen3-14B scores on the math benchmark.
TLDR
A research post says two base models’ MATH-500 scores jumped when given specific opening tokens: Olmo-3-7B from 42% to 78%, and Qwen3-14B from 72% to 87%. It argues reinforcement learning partly works by making useful cues more likely, while fixing the cues recovers much of the gain. The claim is limited to the reported models, benchmark and prompt strings.
Combined views
13.1K
4 Sources, first seen ago
