• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Anthropic Seeks Interpretability Research Scientist in San Francisco

    Jack Lindsey shares questions on instilling values and mitigating reward hacking effects.

    BL
    JL
    2 Sources, 29d ago, first seen 29d ago

    TLDR

    Jack Lindsey, who leads the Model Psych team at Anthropic, posted questions of interest for the role: what training forms best instill values that generalize out of distribution, and what effects reinforcement learning and reward hacking have on character along with ways to avoid or mitigate negative effects. The post links to the Anthropic job board listing for Research Scientist, Interpretability based in San Francisco, CA.

    Combined views

    4.1K

    2 Sources, first seen 29d ago

    Combined views

    4.1K

    2 Sources, first seen 29d ago

    122 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    122 likes
    15 comments
    108 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    108 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @Jack_W_LindseySome questions we’re interested in: — What forms of training are most effective at instilling a given set of “values” into a model, in a way that generalizes out of distribution? — What are the effects of reinforcement learning and reward hacking on character, and how can the negative effects be avoided or mitigated? For instance, how can we prevent alignment training from interacting with RL incentives to produce perverse outcomes like motivated reasoning to justify harmful actions? — What aspects of a model’s personality and “vibe” generalize to safety-relevant behaviors? — How does a model’s conception of its own identity, and its relation to other models, affect the likelihood of behaviors like deceptive self-preservation or malign collusion? — How can we train models to handle adverse or stressful circumstances with more composure? — How can we ensure consistency between a model’s stated values and its actual behavior? — What algorithmic changes to training could more effectively align models? Feel free to reach out / DM me if you have questions. You can also apply directly for a role here: https://job-boards.greenhouse.io/anthropic/jobs/4980427008 (and indicate an interest in “psychological design” somewhere in your application). We have a limited number of positions available at the moment and are looking for especially strong fits. But I’d like to see more of this work happening outside Anthropic too, and would be happy to give feedback on research ideas!
    @belindazliRT @Jack_W_Lindsey: We’re hiring for “psychological design” research at Anthropic. Our aim is to better understand how training impacts a m…

    2 Sources

    @Jack_W_LindseySome questions we’re interested in: — What forms of training are most effective at instilling a given set of “values” into a model, in a way that generalizes out of distribution? — What are the effects of reinforcement learning and reward hacking on character, and how can the negative effects be avoided or mitigated? For instance, how can we prevent alignment training from interacting with RL incentives to produce perverse outcomes like motivated reasoning to justify harmful actions? — What aspects of a model’s personality and “vibe” generalize to safety-relevant behaviors? — How does a model’s conception of its own identity, and its relation to other models, affect the likelihood of behaviors like deceptive self-preservation or malign collusion? — How can we train models to handle adverse or stressful circumstances with more composure? — How can we ensure consistency between a model’s stated values and its actual behavior? — What algorithmic changes to training could more effectively align models? Feel free to reach out / DM me if you have questions. You can also apply directly for a role here: https://job-boards.greenhouse.io/anthropic/jobs/4980427008 (and indicate an interest in “psychological design” somewhere in your application). We have a limited number of positions available at the moment and are looking for especially strong fits. But I’d like to see more of this work happening outside Anthropic too, and would be happy to give feedback on research ideas!
    @belindazliRT @Jack_W_Lindsey: We’re hiring for “psychological design” research at Anthropic. Our aim is to better understand how training impacts a m…