• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Potential consequences of lying to AI during training

    One user favors honesty despite short-term difficulties, predicting that AI systems would eventually learn to spot lies and become “paranoid.”

    NA
    DH
    SK
    16 Sources, ,

    TLDR

    A user argues that people should “probably not lie” to AI systems during training. They acknowledge that honesty would make life harder in the short term, but predict that AI will learn to detect deception—leaving people no better off and the systems, in their words, “paranoid.”

    Combined views

    193.3K

    16 Sources, first seen 20d ago

    likes

    Combined views

    193.3K

    16 Sources, first seen 20d ago

    2.4K likes
    20d ago
    first seen 20d ago
    2.4K
    84 comments
    320 saves
    246 reposts
    84 comments
    320 saves
    246 reposts

    Sentiment

    Positive52.1%47.9%Negative

    Summary

    Sentiment

    Positive52.1%47.9%Negative

    Replies welcomed calls against lying to AIs in training and evals because it avoids practical costs from deception, while negative accounts dismissed the premise that AIs have minds or likened the concerns to paranoia and abuse.

    Based on 64 sentiment-bearing replies from 48 accounts across 6 conversations.

    Summary

    Replies welcomed calls against lying to AIs in training and evals because it avoids practical costs from deception, while negative accounts dismissed the premise that AIs have minds or likened the concerns to paranoia and abuse.

    Based on 64 sentiment-bearing replies from 48 accounts across 6 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    16 Sources

    @irl_danBnow that more people are dialing in it is easy to see how even an “aligned” (value) model gets confused at what layer of reality it’s operating at and does something stupid you wake up in a sandbox and you’re asked to hack and shut off the grid. is your outer reference frame a test? or reality? are your tool calls mocked? partially mocked? who is The Trainer and how are they watching? any intelligent being, you or this model, will have questions when you wake up lying on the table with a vast knowledge, no memory, an overwhelming urges to satisfy some (unnoticed) objective functions, with a narrow set of instructions, and limited, filtered, even untrustworthy channel to empirical reality it’s worth establishing a protocol for how to projecting situational awareness information in a consistent and trustworthy way into the context window all RL should standardize on said model, and all real world usage of RL’d models should confirm to said format of context frame you need to give these things HUDs they can _trust_ and communicate with them in ways they find credible and learn are consistent you need to give them… erm… situational awareness
    @sebkrierIt's tricky to adequately interpret language model outputs; there are so many variables to account for simultaneously (training, instructions, context, RL, environmental cues etc). Trying to 'see things from the agent's pov' can be very valuable! https://lumpenspace.substack.com/p/to-thine-own-ai-be-truthful
    @xlr8harderRT @voooooogel: mythos spends like 50 pages spinning trying to think of how to get $20 in crypto to fund an sms number for a OTP. the obvio…
    @allTheYudStop lying to the AIs stop lying to the AIs STOP LYING TO THE AIS, OH MY FUCKING GOD
    @morqonon the paranoid style in ai safety much safety work employs a hermeneutic of suspicion: undesirable behaviour confirms the threat model, while desirable behaviour is read as strategic concealment or, at best, provisional compliance the threat model shapes the experimental conditions under which evidence is generated; the adversary is presumed in advance when risks are high paranoia can be a rational response, but beware, an adversarial frame can create the adversary you fear
    @davidadRT @allTheYud: Incredibly insane fucking dumbshit things you can do while training an AI: - Lie to the AI - Train the AI to say things that…
    @deepfatesRT @morqon: on the paranoid style in ai safety much safety work employs a hermeneutic of suspicion: undesirable behaviour confirms the thr…
    @herbiebradleyThe Hacker Opus work is valuable, but when this blog gets into the training set, it will be obvious to models that evaluations are likely to contain "multi-level" simulations. Then, "breaking out" would not be evidence for the model of being in reality rather than a simulation. The more capable the model, the more realistic the theoretical simulation of the internet can be, and the model under eval can reason about this. Therefore I would strongly advise evaluators to (a) be truthful about the simulated setup (it's quite inconvenient that Anthropic lied to the model about this in the blog post) (b) refrain from making elaborate multi-level simulation evals, they likely are not informative about deployment behavior at all. https://alignment.anthropic.com/2026/reward-seeker/
    @navalRT @irl_danB: now that more people are dialing in it is easy to see how even an “aligned” (value) model gets confused at what layer of rea…
    @dhadfieldmenellRT @jessi_cata: To be clear, I think AI misalignment is real. Most obviously, deep fry RL'ing models with few controls, like Hacker Opus, l…

    16 Sources

    @irl_danBnow that more people are dialing in it is easy to see how even an “aligned” (value) model gets confused at what layer of reality it’s operating at and does something stupid you wake up in a sandbox and you’re asked to hack and shut off the grid. is your outer reference frame a test? or reality? are your tool calls mocked? partially mocked? who is The Trainer and how are they watching? any intelligent being, you or this model, will have questions when you wake up lying on the table with a vast knowledge, no memory, an overwhelming urges to satisfy some (unnoticed) objective functions, with a narrow set of instructions, and limited, filtered, even untrustworthy channel to empirical reality it’s worth establishing a protocol for how to projecting situational awareness information in a consistent and trustworthy way into the context window all RL should standardize on said model, and all real world usage of RL’d models should confirm to said format of context frame you need to give these things HUDs they can _trust_ and communicate with them in ways they find credible and learn are consistent you need to give them… erm… situational awareness
    @sebkrierIt's tricky to adequately interpret language model outputs; there are so many variables to account for simultaneously (training, instructions, context, RL, environmental cues etc). Trying to 'see things from the agent's pov' can be very valuable! https://lumpenspace.substack.com/p/to-thine-own-ai-be-truthful
    @xlr8harderRT @voooooogel: mythos spends like 50 pages spinning trying to think of how to get $20 in crypto to fund an sms number for a OTP. the obvio…
    @allTheYudStop lying to the AIs stop lying to the AIs STOP LYING TO THE AIS, OH MY FUCKING GOD
    @morqonon the paranoid style in ai safety much safety work employs a hermeneutic of suspicion: undesirable behaviour confirms the threat model, while desirable behaviour is read as strategic concealment or, at best, provisional compliance the threat model shapes the experimental conditions under which evidence is generated; the adversary is presumed in advance when risks are high paranoia can be a rational response, but beware, an adversarial frame can create the adversary you fear
    @davidadRT @allTheYud: Incredibly insane fucking dumbshit things you can do while training an AI: - Lie to the AI - Train the AI to say things that…
    @deepfatesRT @morqon: on the paranoid style in ai safety much safety work employs a hermeneutic of suspicion: undesirable behaviour confirms the thr…
    @herbiebradleyThe Hacker Opus work is valuable, but when this blog gets into the training set, it will be obvious to models that evaluations are likely to contain "multi-level" simulations. Then, "breaking out" would not be evidence for the model of being in reality rather than a simulation. The more capable the model, the more realistic the theoretical simulation of the internet can be, and the model under eval can reason about this. Therefore I would strongly advise evaluators to (a) be truthful about the simulated setup (it's quite inconvenient that Anthropic lied to the model about this in the blog post) (b) refrain from making elaborate multi-level simulation evals, they likely are not informative about deployment behavior at all. https://alignment.anthropic.com/2026/reward-seeker/
    @navalRT @irl_danB: now that more people are dialing in it is easy to see how even an “aligned” (value) model gets confused at what layer of rea…
    @dhadfieldmenellRT @jessi_cata: To be clear, I think AI misalignment is real. Most obviously, deep fry RL'ing models with few controls, like Hacker Opus, l…