Artificial Analysis announces Intelligence Index v4.3 with new automation benchmark
Artificial Analysis says evaluations with private tasks or answers now carry 45% of the index’s weight, up from 40%, while category weights remain unchanged.
TLDR
Artificial Analysis says Intelligence Index v4.3 upgrades Terminal-Bench from 2.1 to 4.0 and replaces τ³-Banking with AutomationBench-AA, its implementation of Zapier’s business workflow automation benchmark. It says AutomationBench-AA uses a private test set in collaboration with Zapier. The organization describes v4.3 as raising the difficulty of AI-agent coding tasks and broadening the workflows tested, bringing forward some changes planned for v5. Category weights remain Agents 30%, Coding 20%, General 30% and Scientific Reasoning 20%, it says.
Combined views
979.3K
11 Sources, first seen 23d ago