• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Steven Hansen on Eval Specific Harnesses and Goodharting

    Gemini Agents lead at Google DeepMind comments on evaluation harnesses.

    NL
    TL
    SH
    5 Sources, 27d ago, first seen 27d ago

    TLDR

    Steven Hansen, identified in the post as Gemini Agents lead at Google DeepMind, stated that eval specific harnesses are just waiting to get Goodharted. The remark comes from his primary work building and leading frontier AI agent systems. The source line presents the comment as a direct quote attributed to him in a conversation on the platform. No additional posts, replies, or confirmations appear in the packet. The statement stands as his expressed view on the topic without further elaboration or context supplied.

    Combined views

    36.7K

    5 Sources, first seen 27d ago

    Combined views

    36.7K

    5 Sources, first seen 27d ago

    456 likes
    456 likes
    33 comments
    76 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    33 comments
    76 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @ZergylordEval specific harnesses are just waiting to get Goodharted.
    @BrendanFoodyAcademic benchmarks are not the right way to evaluate models
    @natolambert@BrendanFoody I wouldn’t say AA is purely academic. I think Claude and GPT moved away from focusing on it maybe a year ago.
    @tallinzenWhat is an “academic benchmark”?

    5 Sources

    @ZergylordEval specific harnesses are just waiting to get Goodharted.
    @BrendanFoodyAcademic benchmarks are not the right way to evaluate models
    @natolambert@BrendanFoody I wouldn’t say AA is purely academic. I think Claude and GPT moved away from focusing on it maybe a year ago.
    @tallinzenWhat is an “academic benchmark”?