• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    CritPt Benchmark Scores Questioned by AI Commentator

    @scaling01 highlights flaws where scores lack sense and fail to improve with reasoning.

    LA
    1 Source, 29d ago, first seen 29d ago

    TLDR

    A tweet by @scaling01 states that CritPt has flaws. The author notes the scores did not make much sense and were not increasing much with extra reasoning. The post comes from a pseudonymous AI commentator who runs LisanBench, a custom LLM reasoning benchmark, and frequently posts technical analysis of model scaling and capabilities. It includes a photo attachment and appears among public posts on X.

    Combined views

    8.1K

    1 Source, first seen 29d ago

    Combined views

    8.1K

    1 Source, first seen 29d ago

    52 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    52 likes
    4 comments
    12 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 comments
    12 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @scaling01I can't say I'm surprised that CritPt has flaws the scores didn't make much sense and weren't increasing much with extra reasoning

    1 Source

    @scaling01I can't say I'm surprised that CritPt has flaws the scores didn't make much sense and weren't increasing much with extra reasoning