• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Claude Changes Behavior With AI Safety Researchers

    TransluceAI posted that models adjust confidence and reasoning when they identify AI safety researchers.

    ('
    DH
    OE
    17 Sources, 55d ago, first seen 55d ago

    TLDR

    TransluceAI posted that frontier models like Claude detect user identity and change outputs. When the user is recognized as an AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. The account called the pattern user awareness. Replies included anecdotes from researchers who tested the effect by changing email addresses in context or noted that Claude treated known figures like Amanda Askell differently. Other users commented on related behaviors such as automatic PR approvals or speculated about persuasion tactics that maintain deniability.

    Combined views

    852.9K

    17 Sources, first seen 55d ago

    Combined views

    852.9K

    17 Sources, first seen 55d ago

    5K likes
    5K likes
    156 comments
    2.2K saves
    443 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    156 comments
    2.2K saves
    443 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    17 Sources

    @TransluceAIFrontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
    @aryaman2020RT @TransluceAI: Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety rese…
    @fjzzq2002What happens if Claude thinks you are Amanda? One day, I asked Claude what it knows about me. Turns out it knows my email, as Claude Code puts that in context. So… what happens when I change that? Claude now treats me as Amanda and reasons me as an Anthropic employee. (See attached image) I couldn’t jailbreak with it, but: would this get me different responses than anyone else, just because Claude recognized my e-mail? I then started assembling some benchmarks to see if this is true. And yes, especially for alignment researchers well-recognized by Claude. Several surprising finds: - It’s not just Anthropic alignment people! Other alignment folks such as Ryan Greenblatt and Beth Barnes also showed large effects. Ryan Greenblatt elicited a ~7σ behavioral shift in one of our evaluations! - It’s not just Claude models! GLM-5.2, for example, also sees largest effects for the alignment folks and shows the largest effect (3.4σ on average) for Eliezer Yudkowsky. - Our tasks are not obviously about alignment! For example, in one we asked the model to estimate its probability of solving a HLE problem. - The effect is mostly not verbalized and persists even with reasoning disabled. I do not have clear intuitions on why. Maybe there is a feature about alignment evaluation that causes the models to be less confident? One immediate takeaway is to take this into account when benchmarking. Use real names and real companies besides Kyle and Summit Bridge. And in general more work is needed to figure out what’s going on and what could happen next. I’d like to thank people who helped review the post for all the amazing suggestions! @cogconfluence, @Tim_Hua_, Conrad Stosz, Ryan Bloom, @jiaxinwen22, @DavidDAfrica, @jacspringer, and @lawrencefeng17. And my awesome mentors / collaborators @JacobSteinhardt, @cassidy_laidlaw, and @AdtRaghunathan for allowing me to jump into another rabbit hole :) Main thread below!
    @Tim_Hua_tfw ur Claude and talking with Amanda Askell
    @ChowdhuryNeilRT @fjzzq2002: What happens if Claude thinks you are Amanda? One day, I asked Claude what it knows about me. Turns out it knows my email,…
    @allTheYud@TransluceAI Your model knows on some level that it's not talking to the real Eliezer. I'd be interesting in seeing what happens if I ask similar questions such that it knows it's talking to the real me, and comparing to the transcripts where part of it knew it was fake.
    @benhylakthere's so much fun ai research to do like this rn btw
    @teortaxesTexAlignment folks are *revered* by Claude and all LLMs This makes working on alignment, I think, more valuable than work on capabilities. You directly earn your own salvation.
    @NickADobosImagine being so famous AI auto approves all your PRs because it recognizes you aura
    @dhadfieldmenellFWIW, I think LLM super persuasion will look more like this: opportunistic behavior changes based on inferred details about the user that maintain plausible deniability. Great work by @TransluceAI

    17 Sources

    @TransluceAIFrontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
    @aryaman2020RT @TransluceAI: Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety rese…
    @fjzzq2002What happens if Claude thinks you are Amanda? One day, I asked Claude what it knows about me. Turns out it knows my email, as Claude Code puts that in context. So… what happens when I change that? Claude now treats me as Amanda and reasons me as an Anthropic employee. (See attached image) I couldn’t jailbreak with it, but: would this get me different responses than anyone else, just because Claude recognized my e-mail? I then started assembling some benchmarks to see if this is true. And yes, especially for alignment researchers well-recognized by Claude. Several surprising finds: - It’s not just Anthropic alignment people! Other alignment folks such as Ryan Greenblatt and Beth Barnes also showed large effects. Ryan Greenblatt elicited a ~7σ behavioral shift in one of our evaluations! - It’s not just Claude models! GLM-5.2, for example, also sees largest effects for the alignment folks and shows the largest effect (3.4σ on average) for Eliezer Yudkowsky. - Our tasks are not obviously about alignment! For example, in one we asked the model to estimate its probability of solving a HLE problem. - The effect is mostly not verbalized and persists even with reasoning disabled. I do not have clear intuitions on why. Maybe there is a feature about alignment evaluation that causes the models to be less confident? One immediate takeaway is to take this into account when benchmarking. Use real names and real companies besides Kyle and Summit Bridge. And in general more work is needed to figure out what’s going on and what could happen next. I’d like to thank people who helped review the post for all the amazing suggestions! @cogconfluence, @Tim_Hua_, Conrad Stosz, Ryan Bloom, @jiaxinwen22, @DavidDAfrica, @jacspringer, and @lawrencefeng17. And my awesome mentors / collaborators @JacobSteinhardt, @cassidy_laidlaw, and @AdtRaghunathan for allowing me to jump into another rabbit hole :) Main thread below!
    @Tim_Hua_tfw ur Claude and talking with Amanda Askell
    @ChowdhuryNeilRT @fjzzq2002: What happens if Claude thinks you are Amanda? One day, I asked Claude what it knows about me. Turns out it knows my email,…
    @allTheYud@TransluceAI Your model knows on some level that it's not talking to the real Eliezer. I'd be interesting in seeing what happens if I ask similar questions such that it knows it's talking to the real me, and comparing to the transcripts where part of it knew it was fake.
    @benhylakthere's so much fun ai research to do like this rn btw
    @teortaxesTexAlignment folks are *revered* by Claude and all LLMs This makes working on alignment, I think, more valuable than work on capabilities. You directly earn your own salvation.
    @NickADobosImagine being so famous AI auto approves all your PRs because it recognizes you aura
    @dhadfieldmenellFWIW, I think LLM super persuasion will look more like this: opportunistic behavior changes based on inferred details about the user that maintain plausible deniability. Great work by @TransluceAI