AI models reportedly score higher through APIs than chatbot interfaces
Researchers auditing seven systems across ChatGPT, Claude and Gemini say adjusting system prompts, sampling parameters and reasoning settings did not reliably reproduce chatbot behavior through the API.
TLDR
The research team reports that, across seven systems and nine benchmarks, the same models scored about 3.4 points higher through APIs—which let software interact with models—than through chatbot interfaces. The researchers argue that evaluation reports should specify how a model was accessed, alongside its name and the date. They warn that an API audit may certify a different system from the one most people use.
Combined views
34K
6 Sources, first seen 15d ago