• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Reward hacking and the challenge of understanding AI models

    In a conversation with Goodfire AI CEO Eric Ho, the interviewer explores reward hacking and ways to monitor models.

    MT
    1 Source, ,

    TLDR

    The interviewer shares a conversation with Goodfire AI CEO Eric Ho about reward hacking, mechanistic interpretability and model monitoring. In a later post, he argues that models are “grown, not built” and that we don’t really know how they work, making better interpretability urgent.

    Combined views

    3.4K

    1 Source, first seen 11h ago

    Combined views

    3.4K

    1 Source, first seen 11h ago

    5 likes
    11h ago
    first seen 11h ago
    5 likes
    4 comments
    1 saves
    1 reposts
    4 comments
    1 saves
    1 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @mattturckStill one aspect of AI that's most poorly understood outside AI circles... "models are grown, not built" and we don't *really* know how they work. Hence the urgent need for better interpretability -- @eric_ho @GoodfireAI11h

    1 Source

    @mattturckStill one aspect of AI that's most poorly understood outside AI circles... "models are grown, not built" and we don't *really* know how they work. Hence the urgent need for better interpretability -- @eric_ho @GoodfireAI11h