Perplexity is very good at building these benchmarks. Whenever I study them in-depth, it’s always the strongest in the field. DRACO was excellent compared to all the other Deep Research ones, for instance. https://twitter.com/perplexity_ai/status/2077099503723946121
the companies that figure out how to capture their value propositions into evals and envs will win.
the ability to create high quality evaluations is a new type of strategic advantage, and it is more important than access to capital, economies of scale, network effects, or any…
The Perplexity is very good at building these benchmarks. Whenever I study them in-depth, it’s always the strongest in the field. DRACO was excellent compared to all the other Deep Research ones, for instance. https://twitter.com/perplexity_ai/status/2077099503723946121
Perplexity has the best (both on cost and performance) deep and wide research harness in Computer. One of the contributing factors is strong internal evals and benchmarks. Today, we're open-sourcing WANDR, the benchmark we use internally for measuring research capabilities.…