The Intelligence Index's shift from four exam-style tests to 10 evaluations
In its September 26, 2026 retrospective, Artificial Analysis says all frontier models use reasoning tokens to “think” before answering, two years after its coverage of OpenAI’s o1-preview.
TLDR
Artificial Analysis describes two years of change since its coverage of OpenAI’s o1-preview. As of September 26, 2026, it says all frontier models use reasoning tokens before answering. Its Intelligence Index has changed too: version 1 used four single-turn, exam-style evaluations covering general knowledge, science, math and basic coding. Version 4.3 incorporates 10 difficult evaluations, including long-horizon tasks for AI agents, challenging coding problems and knowledge work.