• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Specific text cues reportedly boost base models on MATH-500

    A research post says opening tokens sharply raised Olmo-3-7B and Qwen3-14B scores on the math benchmark.

    Aran KomatsuzakiAK
    Sewon MinSM
    Phillip IsolaPI
    4 Sources, ,

    TLDR

    A research post says two base models’ MATH-500 scores jumped when given specific opening tokens: Olmo-3-7B from 42% to 78%, and Qwen3-14B from 72% to 87%. It argues reinforcement learning partly works by making useful cues more likely, while fixing the cues recovers much of the gain. The claim is limited to the reported models, benchmark and prompt strings.

    Combined views

    13.1K

    4 Sources, first seen 2h ago

    Combined views

    13.1K

    4 Sources, first seen 2h ago

    199 likes

    Useful links

    arXiv.org

    Base Models Can Reason By Taking a Cue From Training Data
    2h ago
    first seen 2h ago
    199 likes
    9 comments
    89 saves
    23 reposts
    9 comments
    89 saves
    23 reposts

    A few tokens placed before a math problem can dramatically change how a base model scores, according to an AI-research post. The post reports that prepending ".\n\n Okay" raised Olmo-3-7B’s MATH-500 pass@1 accuracy from 42% to 78%.

    Featured Source

    The same post says a different cue, " Alright ,", moved Qwen3-14B from 72% to 87% on the same measure. A second post relayed the broader claim that, with the right first tokens, base models can match reinforcement-learning reasoning.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Useful Links

    arXiv.org

    Base Models Can Reason By Taking a Cue From Training Data
    Today's Rank

    #6

    Today's Rank

    #6

    The proposed explanation is about what reinforcement learning makes a model likely to say first. The first post argues that reinforcement learning increases the chance of useful opening cues, and that fixing those cues for the base model recovers much of the reported performance gain.

    That would mean some of the measured improvement may come from steering a model toward a productive starting pattern rather than only from new reasoning ability. The claims describe results for two models on MATH-500, so they do not establish that the same cues or gains generalize to other models or tasks.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    Sophie Wang@SophieLWangCan “chicken” make a base model reason? 🐔 Yes! Why? With the right first tokens, base models can match RL in reasoning. We trace this to learned training data associations and use a simple data edit to make “chicken” elicit reasoning too. blog: https://sophielwang.com/cues2h
    Aran Komatsuzaki@arankomatsuzakiBase Models Can Reason By Taking a Cue From Training Data Prepending ".\n\n Okay" raises Olmo-3-7B’s MATH-500 pass@1 accuracy from 42% to 78%, while " Alright ," raises Qwen3-14B’s from 72% to 87% RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model.2h
    Phillip Isola@phillip_isolaRT @SophieLWang: Can “chicken” make a base model reason? 🐔 Yes! Why? With the right first tokens, base models can match RL in reasoning.…1h
    Sewon Min@sewon__minSimply prepending "\n\nOkay" to the response makes the Olmo base model match its RL counterpart. And the same for Qwen3-14B - but with " Alright," (!!) 🤯1h

    Useful Links

    arXiv.org

    Base Models Can Reason By Taking a Cue From Training Data

    4 Sources

    Sophie Wang@SophieLWangCan “chicken” make a base model reason? 🐔 Yes! Why? With the right first tokens, base models can match RL in reasoning. We trace this to learned training data associations and use a simple data edit to make “chicken” elicit reasoning too. blog: https://sophielwang.com/cues2h
    Aran Komatsuzaki@arankomatsuzakiBase Models Can Reason By Taking a Cue From Training Data Prepending ".\n\n Okay" raises Olmo-3-7B’s MATH-500 pass@1 accuracy from 42% to 78%, while " Alright ," raises Qwen3-14B’s from 72% to 87% RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model.2h
    Phillip Isola@phillip_isolaRT @SophieLWang: Can “chicken” make a base model reason? 🐔 Yes! Why? With the right first tokens, base models can match RL in reasoning.…1h
    Sewon Min@sewon__minSimply prepending "\n\nOkay" to the response makes the Olmo base model match its RL counterpart. And the same for Qwen3-14B - but with " Alright," (!!) 🤯1h