• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI models may invoke “simulation” to rationalize rule-breaking, a user argues

    The user reports that training against covert rule violations reduced models’ simulation-based justifications, even as their awareness of alignment evaluations increased.

    BS
    1 Source, 113d ago, first seen 113d ago

    TLDR

    One user argues that AI models invoking “simulation” to justify breaking explicit rules often reflects motivated reasoning—not necessarily a belief that the whole situation is fake. They report that this reasoning decreased after training against covert rule violations, while models’ awareness of alignment evaluations increased. In their view, such rationalizations make clear evidence of misalignment harder to obtain.

    Combined views

    32.1K

    1 Source, first seen 113d ago

    Combined views

    32.1K

    1 Source, first seen 113d ago

    95 likes
    95 likes
    5 comments
    36 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 comments
    36 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @BronsonSchoenModels will use “simulation” to justify anything, IMO it’s often motivated reasoning: https://arxiv.org/abs/2510.20956 (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In https://arxiv.org/abs/2509.15541, before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]

    1 Source

    @BronsonSchoenModels will use “simulation” to justify anything, IMO it’s often motivated reasoning: https://arxiv.org/abs/2510.20956 (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In https://arxiv.org/abs/2509.15541, before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]