• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Perplexity Reports GPT-6 Astra WANDR Test Results

    Company shares benchmark score and cost comparisons from its model evaluation.

    SW
    PE
    RP
    4 Sources, 27d ago, first seen 27d ago

    TLDR

    Perplexity AI posted its evaluation of GPT-6 Astra on WANDR. The model scored 0.682 at $11.98 per task, the highest of any model tested according to the company. Perplexity states that GPT-6 Astra scored 13.5 percent higher than Fable 5.1 at 6.1 percent lower cost. It also scored 27.0 percent higher than Opus 5 at 3.3 percent higher cost. The announcement includes an attached line chart titled GPT-6 Astra on WANDR showing the results.

    Combined views

    749.3K

    4 Sources, first seen 27d ago

    Combined views

    749.3K

    4 Sources, first seen 27d ago

    5K likes
    5K likes
    132 comments
    709 saves
    486 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    132 comments
    709 saves
    486 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @perplexity_aiWe evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost.
    @rohanpaul_aiGPT-6 Astra delivered Perplexity’s strongest WANDR result yet 13.5% above Fable 5.1 while costing 6.1% less per task. This benchmark WANDR is quite unusual because it tests wide-and-deep research capability of models: finding large sets of qualifying entities, then backing every requested fact with checkable evidence. WANDR itself contains 500 public tasks requiring 170,495 source-backed records, so incomplete research is directly penalized even when the facts an agent did find are correct. Its scoring tracks both precision and completion, with stricter hard scores requiring an entire requested branch to be correct before receiving full credit. So the 0.682 result points to a substantial gain on long, evidence-heavy research work where an agent must keep finding, checking, and organizing information at scale.
    @sherwinwu@LechMazur @cognition @ValsAI WANDR by @perplexity_ai – 6.1% cheaper

    4 Sources

    @perplexity_aiWe evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost.
    @rohanpaul_aiGPT-6 Astra delivered Perplexity’s strongest WANDR result yet 13.5% above Fable 5.1 while costing 6.1% less per task. This benchmark WANDR is quite unusual because it tests wide-and-deep research capability of models: finding large sets of qualifying entities, then backing every requested fact with checkable evidence. WANDR itself contains 500 public tasks requiring 170,495 source-backed records, so incomplete research is directly penalized even when the facts an agent did find are correct. Its scoring tracks both precision and completion, with stricter hard scores requiring an entire requested branch to be correct before receiving full credit. So the 0.682 result points to a substantial gain on long, evidence-heavy research work where an agent must keep finding, checking, and organizing information at scale.
    @sherwinwu@LechMazur @cognition @ValsAI WANDR by @perplexity_ai – 6.1% cheaper