The case for judging AI models on separate “smart” and “dumb” axes
One user argues that impressive knowledge and benchmark performance can coexist with poor behavior on work tasks—and that “smart” and “dumb” should be assessed separately.
TLDR
One post argues that an AI model can be highly capable and still make seemingly senseless mistakes. The author points to Gemini models’ knowledge and benchmark capabilities while criticizing their behavior when asked to do work. In another comparison, the author calls Fable 5.1 not quite as “smart” as Astra but significantly less “dumb.” A reply says this framing has helped, to some extent, with using and testing language models.
Combined views
681
1 Source, first seen 2h ago