• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Reasoning Models Struggle on Human-Like Problems

    Tweet shares arXiv paper findings on AI reasoning matching human difficulties.

    RP
    1 Source, 25d ago, first seen 25d ago

    TLDR

    Rohan Paul posted about an arXiv paper examining reasoning models. The post states that these models encounter the same problems that slow human thinking. It notes reasoning token counts can signal problem difficulty yet cautions against relying on a single model run for that signal. The researchers compared human response times to model behavior as part of the work. The tweet included a screenshot of the paper's first page.

    Combined views

    6.3K

    1 Source, first seen 25d ago

    Combined views

    6.3K

    1 Source, first seen 25d ago

    68 likes
    68 likes
    10 comments
    34 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 comments
    34 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @rohanpaul_aiReasoning models struggle on many of the same problems humans do, The same problems often slow down humans and reasoning models. Reasoning-token count can tell you something about how hard a problem is, but this paper says not to trust the result from just 1 model run. The researchers compared human response time with how many reasoning tokens models used on 160 commonsense “best explanation” problems. Questions that took humans longer generally made models think longer too, and humans and models also tended to get the same questions wrong. The interesting part is that this signal became clearer when models tried different reasoning paths and the results were averaged. For GPT-OSS-20B, the human-model correlation rose from 0.41 to 0.55. So chain-of-thought length can be a useful rough signal of problem difficulty, but measure it across multiple runs. That makes model effort more useful as an evaluation signal.

    1 Source

    @rohanpaul_aiReasoning models struggle on many of the same problems humans do, The same problems often slow down humans and reasoning models. Reasoning-token count can tell you something about how hard a problem is, but this paper says not to trust the result from just 1 model run. The researchers compared human response time with how many reasoning tokens models used on 160 commonsense “best explanation” problems. Questions that took humans longer generally made models think longer too, and humans and models also tended to get the same questions wrong. The interesting part is that this signal became clearer when models tried different reasoning paths and the results were averaged. For GPT-OSS-20B, the human-model correlation rose from 0.41 to 0.55. So chain-of-thought length can be a useful rough signal of problem difficulty, but measure it across multiple runs. That makes model effort more useful as an evaluation signal.