Two API Settings Triple GPT-5.6 ARC-AGI-3 Scores
OpenAI engineers show retained reasoning and compaction lift base model results on complex tasks.
TLDR
OpenAI researchers demonstrated that enabling multi-context reasoning and canonical compaction in the API harness allows GPT-5.6 to retain its own thoughts across windows and produce more efficient outputs. These two settings, already deployed in ChatGPT and Codex, transform the model's results on the ARC-AGI-3 public benchmark and raise token efficiency. The findings underscore that benchmark performance reflects both model capability and product harness rather than the model in isolation. A supporting blog post includes animations illustrating the gains in reasoning coherence.
Combined views
5.8M
55 Sources, first seen ago
