• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Minervini Says Model Faithfulness Depends Heavily on Environment

    University of Edinburgh lecturer comments on aryopg et al research into model faithfulness.

    PM
    AP
    CL
    7 Sources, 30d ago, first seen 30d ago

    TLDR

    Pasquale Minervini, tenured Lecturer in NLP/ML at the University of Edinburgh and affiliated with EdinburghNLP, posted that whether a model is faithful depends heavily on environmental factors. He stated that faithfulness, mechinterp, and safety research should take those factors into account. Minervini called the work by @aryopg et al amazing and linked to their post at twitter.com/aryopg/status/2094727222926155985. He is also Co-Founder and CTO of Miniml.AI.

    Combined views

    156.5K

    7 Sources, first seen 30d ago

    Combined views

    156.5K

    7 Sources, first seen 30d ago

    117 likes
    117 likes
    3 comments
    41 saves
    19 reposts

    Sentiment

    Positiveโ€”โ€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    41 saves
    19 reposts

    Sentiment

    Positiveโ€”โ€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    โ€”

    Not ranked yet

    Today's Rank

    โ€”

    Not ranked yet

    7 Sources

    @aryopgCoT monitoring assumes reasoning traces record what shapes an answer. Yet, we found that models are less likely to verbalize a cue in its CoT when the cue comes from a tool return instead of the user message! ๐Ÿ‘€ ๐Ÿงต
    @PMinerviniWhether a model is faithful (and, e.g., whether it's relying on a cue) depends heavily on environmental factors, and faithfulness/mechinterp/safety research should take that into account! ๐Ÿš€ Amazing work by @aryopg et al.
    @Chenyang_Lyuthis is interesting, how reasoning trace leads to answer should take latent states into consideration

    7 Sources

    @aryopgCoT monitoring assumes reasoning traces record what shapes an answer. Yet, we found that models are less likely to verbalize a cue in its CoT when the cue comes from a tool return instead of the user message! ๐Ÿ‘€ ๐Ÿงต
    @PMinerviniWhether a model is faithful (and, e.g., whether it's relying on a cue) depends heavily on environmental factors, and faithfulness/mechinterp/safety research should take that into account! ๐Ÿš€ Amazing work by @aryopg et al.
    @Chenyang_Lyuthis is interesting, how reasoning trace leads to answer should take latent states into consideration