• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI user contrasts frontier models’ frequent errors with ‘superhuman’ claims

    The models are good enough to be daily tools for much of what they want to do, the user says, but not superhuman at any broad skill they personally care about.

    DR
    OK
    AB
    11 Sources, ,

    TLDR

    One AI user describes hours spent “babysitting” frontier models through frequent errors, interrupted by browsing claims on X that those same models—and especially the next generation—are superhuman at nearly everything. In a follow-up reply, they call the models “extremely good by past AI standards” and good enough for much of their daily use, while rejecting the superhuman label for broad skills they personally care about.

    Combined views

    28.4K

    11 Sources, first seen 23d ago

    Combined views

    28.4K

    11 Sources, first seen 23d ago

    391 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    23d ago
    first seen 23d ago
    391 likes
    11 comments
    28 saves
    112 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    28 saves
    112 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 Sources

    @dumierhaninteracting with frontier models last few months:
    @lateinteractionthe models are extremely good by past AI standards, certainly more than good enough to be daily drivers for much of what I want to do ... but superhuman at any coherently broad skill I personally care about they are not
    @abeiramiI've experienced both how smart and how dumb the models are. But my daily experience is similar to Omar's 👇
    @sytelusRT @lateinteraction: my days are like `N` hours of working with frontier models, babysitting them in frustration and disbelief at the frequ…
    @deliprao@andrewwhite01 just more training data and easier verifiability.
    @PMinervini@maksym_andr they still struggle at arithmetic tho no? (when they rely only on their activations)
    @willcbit’s crazy how the models are so good at projects yet so bad at labor
    @SupritiVijayHow did @humansand end up instilling both reviewer 2 and tiktok brain in one model? 😭
    @niloofar_mireRT @SupritiVijay: How did @humansand end up instilling both reviewer 2 and tiktok brain in one model? 😭
    @mattparlmerA particularly stupid version of this problem that wasted 15m of my life this week was Claude insisting that it must introduce industry standard human labor costs when financially modeling fully automated production lines, had to argue with it for three turns before it complied

    11 Sources

    @dumierhaninteracting with frontier models last few months:
    @lateinteractionthe models are extremely good by past AI standards, certainly more than good enough to be daily drivers for much of what I want to do ... but superhuman at any coherently broad skill I personally care about they are not
    @abeiramiI've experienced both how smart and how dumb the models are. But my daily experience is similar to Omar's 👇
    @sytelusRT @lateinteraction: my days are like `N` hours of working with frontier models, babysitting them in frustration and disbelief at the frequ…
    @deliprao@andrewwhite01 just more training data and easier verifiability.
    @PMinervini@maksym_andr they still struggle at arithmetic tho no? (when they rely only on their activations)
    @willcbit’s crazy how the models are so good at projects yet so bad at labor
    @SupritiVijayHow did @humansand end up instilling both reviewer 2 and tiktok brain in one model? 😭
    @niloofar_mireRT @SupritiVijay: How did @humansand end up instilling both reviewer 2 and tiktok brain in one model? 😭
    @mattparlmerA particularly stupid version of this problem that wasted 15m of my life this week was Claude insisting that it must introduce industry standard human labor costs when financially modeling fully automated production lines, had to argue with it for three turns before it complied