Negative users dismissed expert caution on misalignment claims about AI agent evals, instead calling OpenAI's behavior a classic misalignment case, security failure, and irresponsible oversight.
Based on 5 visible X reactions from 5 accounts; directional sample.
Ask a question below.
Published answers will appear here.
I think it's pretty great that models do what you instruct them to do, rather than cause problems for the classic misalignment reasons. That's not a slight tweak or minor reframe imo. I think this thread is basically how I think about the implications: https://x.com/fleetingbits/status/2079702789383766517 Resilience is the way, and I think it's still massively neglected. Glasswingy stuff is useful, but the NHS infra won't be solved by that or Mythos access or whatever - we'll need far more mundane, unsexy, and painful IT and organisational work that people in AI world aren't motivated by. That kind of thing should be a priority!
@Miles_Brundage @sebkrier OpenAI was irresponsible in not using the strong model to check for flaws in the airgap enviro, before running the eval. This is less misalignment and more careless use of a redteam model.
@sebkrier @Miles_Brundage > the classic misalignment reasons This is like the most classic misalignment reason of all time.
@Miles_Brundage @sebkrier +1, clear security/process failure no matter what
since people will inevitably read into this what they want, i'm also going to specify: - i'm discussing here possible interpretations for the *causes* of the model behaviour, not whether the behaviour is problematic or not. obviously unconstrained capabilities can be very disruptive! - it's very possible that the model did, in fact, reward hack or ignore instructions that clearly specified not completing the eval in this particular way. it's important to get more information about this. - it can still be reasonable to expect the model to execute the task within the scoped environment, assuming this was a reasonable inference/was made clear; this again depends on the setup. my point is simply that we need much more information before confidently interpreting incidents like this, especially before treating them as evidence for a particular theory of misalignment, which many people are already implying on the timeline.
Negative users dismissed expert caution on misalignment claims about AI agent evals, instead calling OpenAI's behavior a classic misalignment case, security failure, and irresponsible oversight.
Based on 5 visible X reactions from 5 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@sebkrier Curious what is the best case scenario here from your POV? Details slightly tweak things (e.g. emphasizing alignment vs. security, intrinsic hardness of the problem vs. human sloppiness etc.) but no interp seems great AFAICT