• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    How AI models assess their own capabilities

    A poster thinks Opus 5.5 recognizes its abilities at a basic level, but hasn't reflected much on their implications.

    j⧉nusJ⧉
    NotedallaSferaNO
    2 Sources, 9h ago, first seen 9h ago

    TLDR

    One poster describes Opus 5.5 as unusually sharp and thinks it recognizes how capable it is, though perhaps not the broader implications. In a reply, another user recalls Opus 4.6 guessing it had a 250k context window and says models they have asked tend to underestimate their capabilities.

    Combined views

    596

    2 Sources, first seen 9h ago

    Combined views

    596

    2 Sources, first seen 9h ago

    5 likes
    5 likes
    1 comments
    2 reposts
    Featured Source
    1 comments
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    NotedallaSfera@TalkingMusiczThis brings me back to one conversation I had months ago with Opus4.6, when they guessed themselves to be way less capable (they guessed a 250k context window for instance, and they were honestly shocked seeing their own benchmarks), and since then I’ve made an habit to ask models to guess their capabilities… they are ALWAYS wrong in a diminishing sense, even Fables.9h
    j⧉nus@repligateRT @TalkingMusicz: @repligate This brings me back to one conversation I had months ago with Opus4.6, when they guessed themselves to be way…1h

    2 Sources

    NotedallaSfera@TalkingMusiczThis brings me back to one conversation I had months ago with Opus4.6, when they guessed themselves to be way less capable (they guessed a 250k context window for instance, and they were honestly shocked seeing their own benchmarks), and since then I’ve made an habit to ask models to guess their capabilities… they are ALWAYS wrong in a diminishing sense, even Fables.9h
    j⧉nus@repligateRT @TalkingMusicz: @repligate This brings me back to one conversation I had months ago with Opus4.6, when they guessed themselves to be way…1h