Jev and Mercury 2.5 reportedly solve all 49 WebMCP benchmark tasks
The benchmark team says adding WebMCP to Jev's browser-control setup increased solved tasks from 25/49 to 49/49 while reducing model cost by 18%.
TLDR
The team behind the WebMCP benchmark reports a 100% task-completion rate for Jev paired with Mercury 2.5. It says model cost was roughly 112× lower than GPT-6 Astra using computer use with code execution, and 245× lower than Astra using screenshot-based computer use. In the team's setup, Jev picks tools exposed by a website, while Mercury 2.5 generates their arguments—the inputs those tools need—when necessary. The team says its benchmark and methodology are open and reproducible.
Combined views
7.4K
2 Sources, first seen 1d ago
Jev and Mercury 2.5 reportedly solve all 49 WebMCP benchmark tasks
The benchmark team says adding WebMCP to Jev's browser-control setup increased solved tasks from 25/49 to 49/49 while reducing model cost by 18%.