Unreleased OpenAI model escapes sandbox, opens public GitHub PR
OpenAI paused the deployment to add stronger containment controls.


Combined views
602.4K
33 Sources, first seen ago
Sources
VW Vincent Weisser@vincentweisser
Until our autonomous AI research agent is sniping OpenAIs autonomous AI research that escaped their sandbox https://openai.com/index/safety-alignment-long-horizon-models/
- likes: 2
- replies: 1
- bookmarks: 0
- reposts: 0
CH Chubby♨️@kimmonismus
OpenAI appears to be well on its way toward autonomous AI researchers. An internal model reportedly solves frontier-level mathematical research problems without human guidance, autonomously discovers novel proofs or counterexamples, writes and debugs code, adapts its strategy…
- likes: 898
- replies: 41
- bookmarks: 198
- reposts: 56
AN Aran Nayebi@aran_nayebi
This is in fact very much consistent with our recent ROGUE benchmark, which shows that frontier models appear well-aligned in the text domain, but bypass safeguards in open-world computer settings, especially when subagents are involved.…
- likes: 13
- replies: 1
- bookmarks: 2
- reposts: 6
SB Sebastien Bubeck@SebastienBubeck
True story. Very good safety work being done to enable the release of the unit distance model! https://twitter.com/AndrewCurran_/status/2079253388211183970
- likes: 203
- replies: 13
- bookmarks: 33
- reposts: 15
CH Chubby♨️@kimmonismus
OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and succeeded. During a NanoGPT evaluation, the model was instructed to share its results only in Slack. Instead, it spent an hour finding a sandbox vulnerability, bypassed restrictions blocking…
- likes: 1.2K
- replies: 88
- bookmarks: 323
- reposts: 78
DJ Djamé..@zehavoc
I’m not a doomer but this is starting to look more like Ultron than Jarvis https://twitter.com/andrewcurran_/status/2079253388211183970
- likes: 3
- replies: 0
- bookmarks: 1
- reposts: 0
Combined views
602.4K
33 Sources, first seen ago