• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI

Jev and Mercury 2.5 reportedly solve all 49 WebMCP benchmark tasks

The benchmark team says adding WebMCP to Jev's browser-control setup increased solved tasks from 25/49 to 49/49 while reducing model cost by 18%.

Aditya GroverAG
Volodymyr Kuleshov 🇺🇦VK
2 Sources, 20d ago, first seen 20d ago

TLDR

The team behind the WebMCP benchmark reports a 100% task-completion rate for Jev paired with Mercury 2.5. It says model cost was roughly 112× lower than GPT-6 Astra using computer use with code execution, and 245× lower than Astra using screenshot-based computer use. In the team's setup, Jev picks tools exposed by a website, while Mercury 2.5 generates their arguments—the inputs those tools need—when necessary. The team says its benchmark and methodology are open and reproducible.

Combined views

7.4K

2 Sources, first seen 20d ago

115 likes8 comments19 saves176 reposts

Combined views

7.4K

2 Sources, first seen 20d ago

115 likes8 comments19 saves176 reposts

Sentiment

Positive80.9%19.1%Negative

Summary

Many accounts hailed Jev paired with Mercury 2.5 for near-perfect WebMCP benchmark scores at far lower cost than GPT-6, while negative replies called the comparisons misleading and faulted higher error rates.

Based on 51 sentiment-bearing replies from 47 accounts across 2 conversations.

Sentiment

Positive80.9%19.1%Negative

Summary

Many accounts hailed Jev paired with Mercury 2.5 for near-perfect WebMCP benchmark scores at far lower cost than GPT-6, while negative replies called the comparisons misleading and faulted higher error rates.

Based on 51 sentiment-bearing replies from 47 accounts across 2 conversations.

2 Sources

Aditya Grover@adityagrover_Near-perfect score on the WebMCP benchmark combining Jev with Mercury 2.5! Similar to Mercury 2.5, I suspect Jev itself is a diffusion LLM. Generating structured outputs like JSON in parallel is akin to infilling while sampling from a dLLM.20d
Volodymyr Kuleshov 🇺🇦@volokuleshovRT @0xidanlevin: We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost…20d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    2 Sources

    Aditya Grover@adityagrover_Near-perfect score on the WebMCP benchmark combining Jev with Mercury 2.5! Similar to Mercury 2.5, I suspect Jev itself is a diffusion LLM. Generating structured outputs like JSON in parallel is akin to infilling while sampling from a dLLM.20d
    Volodymyr Kuleshov 🇺🇦@volokuleshovRT @0xidanlevin: We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost…20d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet