Gemini 3.7 Flash Tops AA-AnalystAgent Leaderboard
Google's model ranks first on ArtificialAnlys benchmark for quantitative analysis tasks.
Philipp Schmid posted that Gemini 3.7 Flash reached number one on the new AA-AnalystAgent benchmark from ArtificialAnlys. The evaluation covers 80 real-world quantitative analysis tasks across 14 business and scientific domains including finance, healthcare, hydrology, and government appropriations. The agent operates inside an isolated Python 3.12 sandbox. News from Google stated the model combined reasoning and speed to achieve the highest overall accuracy on the leaderboard for complex data analysis tasks.
Gemini 3.7 Flash just took #1 on @ArtificialAnlys new AA-AnalystAgent. AA-AnalystAgent evaluates against 80 real-world quantitative analysis tasks across 14 business and scientific domains (finance, healthcare, hydrology, government appropriations). The Agent run inside an…
