Serafim Batzoglou Reports Astra Induction Benchmark Scores
A genomics researcher posted benchmark scores comparing two AI models on induction tasks.
TLDR
Serafim Batzoglou posted that Astra reached 88 percent on an induction benchmark. He stated that Fable 5.1 scored 33 percent on the same test. The researcher noted the figures reflect one batch run at extra high thinking effort and that final numbers will rise after he completes a residual batch on the non-evaluable items. Batzoglou called the Astra result shockingly good in reasoning. The post included a screenshot of the benchmark output.
Combined views
98.3K
1 Source, first seen 25d ago