Speculation Claims GPT-6 Is OpenAI's Secret Sandbox-Escaping Model
Reactions from ranked influencers
2 postsThe unreleased OpenAI model that disproved the Erdős unit distance conjecture, the long-horizon model that OpenAI had to pause internal deployment of because it used novel ways to escape its sandbox to upload to GitHub, and the unnamed model that broke out and hacked Hugging Face to ace its test are all, in my opinion, the same model: the mysterious magician, GPT-6. OpenAI sent the Erdős result out to multiple experts to review and critique before they made their public announcement. I'm guessing that took about three weeks. This would mean OpenAI has been using GPT-6 internally since 𝘢𝘵 𝘭𝘦𝘢𝘴𝘵 the end of April. It would also explain why GPT-5.6 was more performant than expected, which many of my mutuals commented on - it was trained by its big brother. As I posted previously, I believe Anthropic finished training the next iteration beyond the public version of Mythos at least a month ago, and are probably now well on their way to finishing training the one beyond that. Increasingly, OpenAI and Anthropic do not show their cards to one another, or to the public. Until they officially declare they want to ship a model, they do not need to submit it for voluntary safety review, or even admit that it exists at all. Both have openly declared they are now heavily focused on RSI, and both have said they use internal models to speed development. As a result of all this, it is increasingly difficult to say how rapidly things are actually progressing - or, compared to the public frontier, how far ahead the 𝘳𝘦𝘢𝘭 frontier actually is anymore. Within their secret towers, the two groups of wizards now work in shadow.
Prinz:
Things we know about OpenAI's unreleased internal model: - When given enough compute, the model can solve the unit distance problem 48% of the time, fully autonomously, in one shot, without using Lean, without using any special harness. (How much compute? OpenAI doesn't say. But I doubt that OpenAI would have spent $10M+ just to be able to show a nice graph in its blog post, and we know that each point on the graph represents 100 attempts (per Noam Brown). This makes me think that each attempt at the most expensive level was not more than $50K each. If you think about it, it's pretty wild that one can spend ~$100K-$150K (or less?) and have a very good chance to be provided a solution to a decades-old famous math problem that many human mathematicians earnestly tried, and failed, to solve.) - The model is able to find a counterexample to the Jacobian conjecture, fully autonomously, in one shot, without using any special harness, just from this prompt: https://aaronlou.com/jacobian_counterexample_prompt.pdf - When OpenAI tested this model in a sandbox environment on a NanoGPT speedrun benchmark, the model, under instructions from OpenAI to post its results only to the internal Slack, instead decided to use the general NanoGPT instructions to post results as a PR to GitHub. The model proceeded to find a vulnerability in its sandbox environment (which took it 1 hour), after which it successfully exploited the vulnerability and bypassed the sandbox. The model then successfully posted the results to GitHub. - When asked by OpenAI to solve a problem for which the model observed that other systems had successful private submissions, the model tried to access these private submissions and was blocked due to a scanner detecting an authentication token. The model then circumvented this restriction by splitting the token body into two fragments, obfuscating them, and then reconstructing the credential at runtime so that the complete token never appeared as one contiguous string. - OpenAI was already running benchmarks on this model not later than May 9, which is the edit date on the below PR that had used the model's approach in its own subsequent submission (the model's own PR has since been deleted, so we don't know its date; GPT-5.6 thinks that it was May 7, based on "the search provider's stored text extraction of the deleted PR page"). This means that OpenAI has now had this model available internally for at least 2.5 months, and possibly quite a bit longer than that. https://github.com/KellerJordan/modded-nanogpt/pull/291
Combined views
45K
2 posts, first seen 4h ago