• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Two API Settings Triple GPT-5.6 ARC-AGI-3 Scores

    OpenAI engineers show retained reasoning and compaction lift base model results on complex tasks.

    OpenAIOP
    Sam AltmanSA
    Noam BrownNB
    55 Sources, 68d ago, first seen 68d ago

    TLDR

    OpenAI researchers demonstrated that enabling multi-context reasoning and canonical compaction in the API harness allows GPT-5.6 to retain its own thoughts across windows and produce more efficient outputs. These two settings, already deployed in ChatGPT and Codex, transform the model's results on the ARC-AGI-3 public benchmark and raise token efficiency. The findings underscore that benchmark performance reflects both model capability and product harness rather than the model in isolation. A supporting blog post includes animations illustrating the gains in reasoning coherence.

    Combined views

    5.8M

    55 Sources, first seen 68d ago

    Combined views

    5.8M

    55 Sources, first seen 68d ago

    35.3K likes
    35.3K likes
    1.8K comments
    7.8K saves
    2.2K reposts
    Featured Source
    TSTed Sanders@sandersted4:05 PM · Jul 29, 2026
    860TECH

    On ARC-AGI-3, GPT-5.6 is dumb as dirt. But it turns out if you turn on two API settings that we use in ChatGPT and Codex, its score on the public set rises ~3x and its token efficiency rises ~6x. Perf is a function of model + product, not just the model.

    OpenAIHow enabling two settings tripled our scores on the ARC-AGI-3 benchmarkHow two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
    1.8K comments
    7.8K saves
    2.2K reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    55 Sources

    Ted Sanders@sanderstedOn ARC-AGI-3, GPT-5.6 is dumb as dirt. But it turns out if you turn on two API settings that we use in ChatGPT and Codex, its score on the public set rises ~3x and its token efficiency rises ~6x. Perf is a function of model + product, not just the model. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/68d
    Tibo@thsottiauxTurns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3. Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/68d
    Vaibhav (VB) Srivastav@reach_vb🤯68d
    Andrew Curran@AndrewCurran_@thsottiaux With caffeine, and without caffeine.68d
    Lisan al Gaib@scaling01bruh68d
    bayes@bayeslord@thsottiaux @tszzl cool68d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexNobody is ahead of OpenAI on post-training68d
    Ethan Mollick@emollickModel + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better)68d
    OpenAI@OpenAIWe implemented the harness with the Responses API and turned on: → Retained reasoning → Context compaction On the public set, GPT-5.6 Sol’s score rose 188% while using 6x fewer output tokens.68d
    Noam Brown@polynoamialAutocompaction is extremely effective68d

    55 Sources

    Ted Sanders@sanderstedOn ARC-AGI-3, GPT-5.6 is dumb as dirt. But it turns out if you turn on two API settings that we use in ChatGPT and Codex, its score on the public set rises ~3x and its token efficiency rises ~6x. Perf is a function of model + product, not just the model. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/68d
    Tibo@thsottiauxTurns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3. Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/68d
    Vaibhav (VB) Srivastav@reach_vb🤯68d
    Andrew Curran@AndrewCurran_@thsottiaux With caffeine, and without caffeine.68d
    Lisan al Gaib@scaling01bruh68d
    bayes@bayeslord@thsottiaux @tszzl cool68d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexNobody is ahead of OpenAI on post-training68d
    Ethan Mollick@emollickModel + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better)68d
    OpenAI@OpenAIWe implemented the harness with the Responses API and turned on: → Retained reasoning → Context compaction On the public set, GPT-5.6 Sol’s score rose 188% while using 6x fewer output tokens.68d
    Noam Brown@polynoamialAutocompaction is extremely effective68d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    OpenAIGPT-5.6 SolTiboChatGPTResponses APICodex

    Related

    OpenAI announces $200 Pro usage changes and a pledge to keep five-hour limit off

    Tibo described half the prior plan’s API-priced usage value, while promising more work from model improvements. Parlmer’s separate limit complaint later appeared resolved.

    Codex Engineer Shares Extended Context Setup for GPT-5.6 Sol

    OpenAI Codex engineer Tibo Sottiaux posted configuration steps to unlock the larger window for ChatGPT accounts.

    Tibo Resets Usage Limits for ChatGPT Work and Codex

    OpenAI Codex lead Tibo resets paid account limits to mark GPT-5.6 Sol release.