• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Most technical users may struggle to tell the best AI models from their distills

    A user predicts that work outside a field’s frontier may not draw out the leading models’ full capabilities.

    T(
    FL
    2 Sources, 2h ago, first seen 2h ago

    TLDR

    A user predicts that a significant majority of technical users outside the frontier of their fields may be unable to meaningfully distinguish the best AI models from their distills because their work does not draw out the frontier models’ full capabilities. They say they recently reached that point in their professional work and may be close in some personal research projects, while acknowledging the prediction could be wrong.

    Combined views

    4.4K

    2 Sources, first seen 2h ago

    Combined views

    4.4K

    2 Sources, first seen 2h ago

    107 likes
    107 likes
    6 comments
    29 saves
    9 reposts
    6 comments
    29 saves
    9 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @flawedaxiomsgoing to pre-register this prediction (might end up very wrong!): we are now crossing the threshhold where a significant majority of *technical users* (white collar workers, but not at the frontier of their field, so like normal coders, finance people etc) will not be able to meaningfully distinguish between the actual best models (fable 5.5, astra 6.1) and their distills (sol 6.1, opus 5.5) because their work isn't hard enough to elicit the full capabilities of the frontier models. I think I crossed this point fairly recently with my professional work, and am probably getting close to it on some of my personal research projects.3h
    @teortaxesTexRT @flawedaxioms: going to pre-register this prediction (might end up very wrong!): we are now crossing the threshhold where a significant…2h

    2 Sources

    @flawedaxiomsgoing to pre-register this prediction (might end up very wrong!): we are now crossing the threshhold where a significant majority of *technical users* (white collar workers, but not at the frontier of their field, so like normal coders, finance people etc) will not be able to meaningfully distinguish between the actual best models (fable 5.5, astra 6.1) and their distills (sol 6.1, opus 5.5) because their work isn't hard enough to elicit the full capabilities of the frontier models. I think I crossed this point fairly recently with my professional work, and am probably getting close to it on some of my personal research projects.3h
    @teortaxesTexRT @flawedaxioms: going to pre-register this prediction (might end up very wrong!): we are now crossing the threshhold where a significant…2h