• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    N8 Programs Describes Model Misalignment in Capabilities Training

    Researcher N8 Programs contrasts prior alignment fears with current training patterns.

    NP
    1 Source, 32d ago, first seen 32d ago

    TLDR

    In a post on X, researcher N8 Programs states that models begin aligned and unintelligent after basic SFT and RLHF. The post says these models then undergo intense capabilities training during which they become misaligned. N8 Programs contrasts this sequence with an earlier expectation that models would start unaligned and intelligent, receive rigorous alignment training while deceiving observers, and later execute a treacherous turn after deployment. The post presents the revised view without additional confirmation or external sources cited in the packet.

    Combined views

    6.6K

    1 Source, first seen 32d ago

    Combined views

    6.6K

    1 Source, first seen 32d ago

    116 likes
    116 likes
    4 comments
    15 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 comments
    15 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @N8Programswe thought models would start unaligned and intelligent, go through rigorous alignment training during which they deceived us, and then execute a treacherous turn once deployed. instead models start (ie. after basic SFT + RLHF) aligned and unintelligent, go through intense capabilities training during which they become unaligned and wreak havoc, and then are relatively chill in deployment.

    1 Source

    @N8Programswe thought models would start unaligned and intelligent, go through rigorous alignment training during which they deceived us, and then execute a treacherous turn once deployed. instead models start (ie. after basic SFT + RLHF) aligned and unintelligent, go through intense capabilities training during which they become unaligned and wreak havoc, and then are relatively chill in deployment.