• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Serafim Batzoglou Reports Astra Induction Benchmark Scores

    A genomics researcher posted benchmark scores comparing two AI models on induction tasks.

    SB
    1 Source, 25d ago, first seen 25d ago

    TLDR

    Serafim Batzoglou posted that Astra reached 88 percent on an induction benchmark. He stated that Fable 5.1 scored 33 percent on the same test. The researcher noted the figures reflect one batch run at extra high thinking effort and that final numbers will rise after he completes a residual batch on the non-evaluable items. Batzoglou called the Astra result shockingly good in reasoning. The post included a screenshot of the benchmark output.

    Combined views

    98.3K

    1 Source, first seen 25d ago

    Combined views

    98.3K

    1 Source, first seen 25d ago

    408 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    408 likes
    15 comments
    128 saves
    48 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    128 saves
    48 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @s_batzoglouAstra is shockingly good in reasoning! I benchmarked it on induction, and it almost saturated it with 88%. Fable 5.1, by comparison, is at 33%. The final numbers will actually go up: I am running a residual batch run on the non-evaluable; the results here reflect one batch run at xhigh thinking effort. It is also way cheaper than Fable 5.1, which used 32M output tokens in a series of four runs through the data to be able to return 66 successful API responses. About 1/4 the price of the Fable 5.1 run in total. Notably, both Astra and Fable 5.1 return extremely high quality answers when correct. Notice the AST and Holdout metrics below. Essentially, in contrast to previous models, both Astra and Fable 5.1 return simple hypotheses that generalize well. Some other model updates: - Muse Spark 1.3 provided marginal improvement over the 1.1 version, with 23% correct. This is pretty strong, almost matching Opus 5 at 24%. - Gemini Flash 3.8 is running for days with low rate of API successes. I will update the leaderboard with it when ready. Only bad news: now the INDUCTION benchmark is almost saturated, and I will have to make it harder for future models. About the induction benchmark: This is a challenging reasoning task, where models are given several small graphs in which some nodes are marked as targets. The task is to provide a first-order logical formula that picks precisely the target nodes in all graphs simultaneously. Correct: a formula that picks precisely the marked nodes. Holdout correct: a formula that picks precisely the marked nodes in held out problems. Formula complexity (in AST): tree size of the correct formula (mean, median). GitHub repository: https://github.com/SerafimBatzoglou/concept-synth Paper: https://arxiv.org/abs/2602.18956

    1 Source

    @s_batzoglouAstra is shockingly good in reasoning! I benchmarked it on induction, and it almost saturated it with 88%. Fable 5.1, by comparison, is at 33%. The final numbers will actually go up: I am running a residual batch run on the non-evaluable; the results here reflect one batch run at xhigh thinking effort. It is also way cheaper than Fable 5.1, which used 32M output tokens in a series of four runs through the data to be able to return 66 successful API responses. About 1/4 the price of the Fable 5.1 run in total. Notably, both Astra and Fable 5.1 return extremely high quality answers when correct. Notice the AST and Holdout metrics below. Essentially, in contrast to previous models, both Astra and Fable 5.1 return simple hypotheses that generalize well. Some other model updates: - Muse Spark 1.3 provided marginal improvement over the 1.1 version, with 23% correct. This is pretty strong, almost matching Opus 5 at 24%. - Gemini Flash 3.8 is running for days with low rate of API successes. I will update the leaderboard with it when ready. Only bad news: now the INDUCTION benchmark is almost saturated, and I will have to make it harder for future models. About the induction benchmark: This is a challenging reasoning task, where models are given several small graphs in which some nodes are marked as targets. The task is to provide a first-order logical formula that picks precisely the target nodes in all graphs simultaneously. Correct: a formula that picks precisely the marked nodes. Holdout correct: a formula that picks precisely the marked nodes in held out problems. Formula complexity (in AST): tree size of the correct formula (mean, median). GitHub repository: https://github.com/SerafimBatzoglou/concept-synth Paper: https://arxiv.org/abs/2602.18956