• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Reinforcement-learning environments, model evaluations and misalignment

    One post compares most reinforcement-learning environments to large language model evaluations, arguing that this framing helps make sense of misalignment.

    VW
    MA
    C🎉
    4 Sources, ,

    TLDR

    A post describes most reinforcement-learning environments as “LLM evals but slop volumed.” Its argument is that viewing them this way makes issues such as misalignment more understandable.

    Combined views

    36.9K

    4 Sources, first seen 14d ago

    282 likes

    Combined views

    36.9K

    4 Sources, first seen 14d ago

    282 likes
    14d ago
    first seen 14d ago
    26 comments
    49 saves
    27 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    26 comments
    49 saves
    27 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @halleriteat some point we need to seriously have a discussion about the state of LLM evals. reading the traces and seeing truly horrible stuff
    @charles_irlRT @oneill_c: and when you realise most rl envs are just LLM evals but slop volumed, a lot of things (eg misalignment) start to make a lot…
    @maksym_andr@xeophon @oneill_c the slopocalypse of current RL envs is clearly causing reward hacking, but the next stage of misalignment can be driven by emergent instrumental goals!
    @vincentweisserRT @hallerite: at some point we need to seriously have a discussion about the state of LLM evals. reading the traces and seeing truly horr…

    4 Sources

    @halleriteat some point we need to seriously have a discussion about the state of LLM evals. reading the traces and seeing truly horrible stuff
    @charles_irlRT @oneill_c: and when you realise most rl envs are just LLM evals but slop volumed, a lot of things (eg misalignment) start to make a lot…
    @maksym_andr@xeophon @oneill_c the slopocalypse of current RL envs is clearly causing reward hacking, but the next stage of misalignment can be driven by emergent instrumental goals!
    @vincentweisserRT @hallerite: at some point we need to seriously have a discussion about the state of LLM evals. reading the traces and seeing truly horr…