• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Faithful LLMs reportedly decide and report using the same layers in one research setting

    A researcher sharing a new COLM paper says unfaithful models did not use the same layers for both tasks.

    Jack LindseyJL
    David Atkinson @ COLMDA
    2 Sources, 2h ago, first seen 2h ago

    TLDR

    A researcher sharing the COLM paper “Identifying Introspection From the Inside” asks whether LLMs know what drives their choices when they describe their decisions. In the researchers’ setting, the post says, faithful models used the same layers to decide and report; unfaithful models did not.

    Combined views

    2.5K

    2 Sources, first seen 2h ago

    Combined views

    2.5K

    2 Sources, first seen 2h ago

    53 likes
    53 likes
    2 comments
    25 saves
    22 reposts
    Featured Source
    2 comments
    25 saves
    22 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    David Atkinson @ COLM@diatkinsonNew COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing? In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵2h
    Jack Lindsey@Jack_W_LindseyRT @diatkinson: New COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what driv…1h

    2 Sources

    David Atkinson @ COLM@diatkinsonNew COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing? In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵2h
    Jack Lindsey@Jack_W_LindseyRT @diatkinson: New COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what driv…1h