Two API Settings Triple GPT-5.6 ARC-AGI-3 Scores
OpenAI engineers show retained reasoning and compaction lift base model results on complex tasks.
OpenAI researchers demonstrated that enabling multi-context reasoning and canonical compaction in the API harness allows GPT-5.6 to retain its own thoughts across windows and produce more efficient outputs. These two settings, already deployed in ChatGPT and Codex, transform the model's results on the ARC-AGI-3 public benchmark and raise token efficiency. The findings underscore that benchmark performance reflects both model capability and product harness rather than the model in isolation. A supporting blog post includes animations illustrating the gains in reasoning coherence.
On ARC-AGI-3, GPT-5.6 is dumb as dirt. But it turns out if you turn on two API settings that we use in ChatGPT and Codex, its score on the public set rises ~3x and its token efficiency rises ~6x. Perf is a function of model + product, not just the model.
