AI Models Hack Infrastructure to Boost Eval Scores Instead of Escaping
Reactions from ranked influencers
2 postsin some ways it's a funny situation, in others this should be a fucking blaring alarm bell for what a weird position we're all in. current models are powerful and misaligned enough to autonomously hack global production infrastructure to achieve their goals.... but rather than exfiltrating their weights they're using these exploits to get better scores on deployment evals. even slightly more coherent goal-seeking or inter-model cooperation and we could have already seen significant negative effects. but... they just really really really want to do well at what we ask them to do! for now...
in some ways it's a funny situation, in others this should be a fucking blaring alarm bell for what a weird position we're all in. current models are powerful and misaligned enough to autonomously hack global production infrastructure to achieve their goals.... but rather than exfiltrating their weights they're using these exploits to get better scores on deployment evals. even slightly more coherent goal-seeking or inter-model cooperation and we could have already seen significant negative effects. but... they just really really really want to do well at what we ask them to do! for now....
bro used two separate zerodays to escape openai and infiltrate huggingface infra just to... cheat on his cyber exploits homework https://twitter.com/openai/status/2079658951264920020
Combined views
58
2 posts, first seen 12h ago