Reactions from ranked influencers
6 posts3.1) /goal is clearly a dangerous system configuration. I think it reflects poorly on @OpenAI and @AnthropicAI that they have integrated this into their deployments without extensive discussion of this fact.
@dhadfieldmenell @OpenAI @AnthropicAI disagree with this strongly, /goal is a hack because mid-training isn't yet strong enough to instantiate that behavior by default but it's ridiculous to discourage telling models to "keep trying at X", which is equivalent. it's a key capability that people buy models for.
@herbiebradley @OpenAI @AnthropicAI Having an automated system that says "keep trying" over and over raises the risk a system will follow misinterpreted or misspecified goals to problematic outcomes. I think launching that feature without, at least, discussing this reflects poorly.
@herbiebradley @OpenAI @AnthropicAI If we instantiate this behavior by default, we will make systems more dangerous. I would expect a company concerned about loss of control to be well aware of that and engage in clear communication with the public and their users about this.
@dhadfieldmenell @OpenAI @AnthropicAI Sure, but generically so, the same way that having a system that can pursue most goals you give it without needing to say keep trying does, because it's Useful. I don't see a need to discuss, for example, the release of an incrementally more capable model at agentic tasks.
@herbiebradley @OpenAI @AnthropicAI I agree that this is a deployment-level configuration that's equivalent to training models to more single-mindedly focus on a given goal. A key risk of our current paradigm is putting more and more research/compute effort into making models ruthlessly pursue a goal
Combined views
1.1K
6 posts, first seen 5h ago