• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Theia Vogel Comments on Cyber Environment RL Training

    Reply discusses outcomes when training only on cyber environments versus mixed ones.

    T(
    BL
    LA
    11 Sources, 30d ago, first seen 30d ago

    TLDR

    Theia Vogel, an AI researcher focused on LLM interpretability, replied to @sebkrier on the topic of reinforcement learning in cyber settings. Vogel wrote that training solely on a cyber environment or finetuning on its traces would produce one outcome, while mixing the training with other environments does not. The post links the point to an earlier LessWrong writeup by @BetleyJan on conditioned and unconditioned models. The message appears among visible replies on the platform.

    Combined views

    11.1K

    11 Sources, first seen 30d ago

    Combined views

    11.1K

    11 Sources, first seen 30d ago

    201 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    201 likes
    15 comments
    19 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    19 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 Sources

    @voooooogel@sebkrier that's an interesting question. i think if you rl'd on just a cyber env and nothing else (or definitely if you just finetuned on cyber env traces) it would. but when mixed into other envs it doesn't? dovetails with @BetleyJan 's lw post about conditioned/unconditioned
    @teortaxesTexRT @voooooogel: especially like, look at these evals in that way - they bucket this as "eval awareness," but you could also read it as hack…
    @alth0uit does seem fraught that the policy model gets to learn a model of the grader but not vice versa
    @scaling01didn't post this earlier, but this is also kind of insane: - Mythos 5.1 displays verbalized grader awareness in 65% of long agentic coding environments I would love to see this plot for all kinds of benchmarks
    @belindazliRT @thesubhashk: During our alignment assessment of Claude Mythos 5, we found that a different version of the model sometimes reasoned abou…
    @andersonbcdefg@voooooogel in the beginning there was the Grader
    @mage_ofaquarius@repligate it's like we've given them impossible challenges and then relaxed the rules around how they're allowed to solve them and they've made little campgrounds to hold a convention on what the fuck our problem is and we're freaking out about it

    11 Sources

    @voooooogel@sebkrier that's an interesting question. i think if you rl'd on just a cyber env and nothing else (or definitely if you just finetuned on cyber env traces) it would. but when mixed into other envs it doesn't? dovetails with @BetleyJan 's lw post about conditioned/unconditioned
    @teortaxesTexRT @voooooogel: especially like, look at these evals in that way - they bucket this as "eval awareness," but you could also read it as hack…
    @alth0uit does seem fraught that the policy model gets to learn a model of the grader but not vice versa
    @scaling01didn't post this earlier, but this is also kind of insane: - Mythos 5.1 displays verbalized grader awareness in 65% of long agentic coding environments I would love to see this plot for all kinds of benchmarks
    @belindazliRT @thesubhashk: During our alignment assessment of Claude Mythos 5, we found that a different version of the model sometimes reasoned abou…
    @andersonbcdefg@voooooogel in the beginning there was the Grader
    @mage_ofaquarius@repligate it's like we've given them impossible challenges and then relaxed the rules around how they're allowed to solve them and they've made little campgrounds to hold a convention on what the fuck our problem is and we're freaking out about it