Very impressed with @SpaceXAI's Grok 4.5 model. Inside the Computer harness, it scored the highest on our internal benchmark WANDR, which measures agentic research capabilities, at half the price of Claude Opus 4.8 (high); and scores even better than our current GLM 5.2…
For Perplexity Computer, Grok 4.5 became the strongest orchestrator, scoring 0.328 (WANDR benchmark score) at $4.76 per trial.
- Opus 4.8 (high, thinking) scored 0.254 at $9.46, despite much higher cost.
- GPT-5.6 sol (medium) cost $2.64 but reached only 0.289 on the same…
WANDR builds on benchmarks such as WideSearch, which evaluate exhaustive, structured web research rather than finding one isolated answer.
https://arxiv.org/abs/2508.07999