Anthropic Seeks Interpretability Research Scientist in San Francisco
Jack Lindsey shares questions on instilling values and mitigating reward hacking effects.
Jack Lindsey, who leads the Model Psych team at Anthropic, posted questions of interest for the role: what training forms best instill values that generalize out of distribution, and what effects reinforcement learning and reward hacking have on character along with ways to avoid or mitigate negative effects. The post links to the Anthropic job board listing for Research Scientist, Interpretability based in San Francisco, CA.
Combined views
3.7K
2 posts, first seen 17h ago
