Users offer theories about Astra’s apparent reasoning gains
One user suspects the gains without chain-of-thought are somewhat underestimated, citing saturated benchmarks and saying Astra seems to benefit much more from filler tokens.
TLDR
One user guesses that the numbers somewhat underestimate Astra’s reasoning jump without chain-of-thought—the step-by-step reasoning text a model produces. They argue that ECI struggles with large jumps when benchmarks are saturated, and note that the test they cite omitted filler tokens, which Astra seems to benefit from much more. A reply offers a different hypothesis: looping transformers could reuse previously cached outputs to provide a “virtual” chain of thought.
Combined views
118
1 Source, first seen 20d ago