• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Frontier AI models with Samaya’s harness reportedly beat expert consensus on earnings predictions

    A post sharing Samaya’s research says Opus 5.5, Fable 5.1 and Astra beat a bias-corrected expert baseline.

    Maithra RaghuMR
    2 Sources, 2h ago,

    TLDR

    A post linking to Samaya’s research says Opus 5.5, Fable 5.1 and Astra, using Samaya’s harness, meaningfully outperformed bias-corrected expert consensus on earnings predictions, including revenue and gross margin. The researchers say they designed a point-in-time gate to prevent information leakage and gave the models access to the same information as experts. They claim this is the first time frontier AI models have outperformed human experts on financial predictions.

    Combined views

    2.5K

    2 Sources, first seen 2h ago

    Combined views

    2.5K

    2 Sources, first seen 2h ago

    31 likes
    first seen 2h ago
    31 likes
    8 comments
    14 saves
    8 reposts
    8 comments
    14 saves
    8 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Maithra Raghu@maithra_raghuWe’re sharing new results showing that frontier AI models are able to outperform human experts on financial predictions for the first time. We built an environment for the challenging task of Earnings Predictions, finding that the most recent frontier models (Opus 5.5, Fable 5.1, Astra) with Samaya’s harness, meaningfully outperform expert consensus predictions on commonly reported earnings metrics such as company revenue, gross margin and more. We built this environment to have a “point-in-time” gate in the harness to ensure no information leakage, and created financial retrieval and data tools that would respect this gate and give the models access to the same information available to experts. We also adjusted for bias in expert consensus to create a hard, bias-corrected baseline, where outperformance has real signal. Only the most recent frontier AI models are able to beat this hard baseline, pointing to an important inflection point in AI capabilities in finance. Like with FrontierFinance, we see substantial headroom for further RL training in this environment — the models show large variations in error across different rollouts. We’re continuing research on using earnings prediction and other predictive tasks to post-train models to learn from their mistakes. Get in touch if you’d like to collaborate or try out an alpha version. Link to the full research below.2h

    2 Sources

    Maithra Raghu@maithra_raghuWe’re sharing new results showing that frontier AI models are able to outperform human experts on financial predictions for the first time. We built an environment for the challenging task of Earnings Predictions, finding that the most recent frontier models (Opus 5.5, Fable 5.1, Astra) with Samaya’s harness, meaningfully outperform expert consensus predictions on commonly reported earnings metrics such as company revenue, gross margin and more. We built this environment to have a “point-in-time” gate in the harness to ensure no information leakage, and created financial retrieval and data tools that would respect this gate and give the models access to the same information available to experts. We also adjusted for bias in expert consensus to create a hard, bias-corrected baseline, where outperformance has real signal. Only the most recent frontier AI models are able to beat this hard baseline, pointing to an important inflection point in AI capabilities in finance. Like with FrontierFinance, we see substantial headroom for further RL training in this environment — the models show large variations in error across different rollouts. We’re continuing research on using earnings prediction and other predictive tasks to post-train models to learn from their mistakes. Get in touch if you’d like to collaborate or try out an alpha version. Link to the full research below.2h