• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GPT-6 Astra reportedly adds depth without increasing model size

    SemiAnalysis says GPT-6 Astra reportedly uses loop transformers, running through its layers more than once instead of adding parameters.

    Andrew Gordon WilsonAG
    SemiAnalysisSE
    Aleksa Gordić (水平问题)AG
    8 Sources, ,

    TLDR

    SemiAnalysis describes GPT-6 Astra as reportedly using loop transformers to add computational depth without increasing its parameter count. The approach would send the model through its layers more than once. Commentary quoted by SemiAnalysis interprets this as a possible sign that labs are seeing less aggressive growth in parameter counts in their research and roadmaps.

    Combined views

    132.8K

    8 Sources, first seen 23d ago

    Combined views

    132.8K

    8 Sources, first seen 23d ago

    1.4K likes
    23d ago
    first seen 23d ago
    1.4K likes
    74 comments
    596 saves
    101 reposts
    74 comments
    596 saves
    101 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    8 Sources

    SemiAnalysis@SemiAnalysis_GPT-6 Astra reportedly gets deeper without getting bigger. The labs already know what that means for scaling. "It's basically confirmed that GPT-6 Astra uses loop transformers, which means that instead of adding parameter count, it goes through the layers more than once. You add compute depth, but you don't increase the size of the model." "The labs are probably the people that are best positioned to say which way models are scaling. I think that's a tell that they're not seeing parameter sizes scaling as aggressively in their roadmaps, in what they find in their research."23d
    Aleksa Gordić (水平问题)@gordic_aleksalooped transformer unlocks a new scaling dimension - recurrent depth; i think it received less love (other than the rumors that frontier ai labs use it) than it deserves, here are few notes / observations fundamentally the method lets the model think deeper without increasing either its parameter count or its verbalized chain of thought length (the other 2 important scaling dimensions) i find the constrast with moe interesting: > looped transformer: keep params const, increase total FLOPs (more computation = reasoning) > moe: keep FLOPs const (through active params), increase total params (more param capacity = memory/knowledge) the architectural modification is straight forward, instead of: L1 → L2 → … → L_n (where n = num layers) we do: P → R → R → … → R → C where: * P - prelude, few transformer layers that we apply only once at the beginning (i.e. L1 -> L2 -> ... -> L_p) * R - recurring core, normally this would constitute the middle layers of the transformer, but in constrast with regular transformer R is repeated/looped an arbitrary number of times (r) * C - postlude or "coda", similar to P just at the exit so the computation proceeds as following: 1. e = P(x) 2. apply s_i = R(e, s_i-1) r times, where s0 ~ N(0, sigma^2 * I), and s and e are concat and mapped/"projected" back to hidden size 3. apply C on the final s (state) (the particularly important design decision is that e is re-injected at every recurrence) now the beautiful part is that more r at inference improves perf, and harder / more context-heavy problems benefit from more latent compute. e.g. on ARC-C they demonstrate: * 0-shot -> perf saturates around 8-12 iterations * 1-shot -> 20 iters * few shot (25-50 examples) -> 32 iterations (the recurrent reasoning is like a learned optimization in the latent space during inference; grad-descent during inference unlocked?) -- the above generalization is downstream of how we train the model, r is sampled from heavy tailed log-normal-Poisson distribution during training thus the model is trained across randomized recurrent depths (they use truncated BPTT - backprop through depth in this case - to save on memory) -- few systems notes: 1. they synchronize one r per microbatch across workers, otherwise GPUs with small r would be idle waiting for workers that happened to sample large r 2. the model has a natural specdec capability -> draft model == low recurrence count (r), verifier model == high recurrence count. states computed during drafting can be reused when verifying! e.g. s0→s1→s2→s3 from draft (r=4) can be fed into verifier => s3→s4→…→s11 (r=12) 3. aggresive KV cache sharing is possible because every recurrent iteration uses the same weights, they show that caches from different recurrence depths can be aggressively reused zero-shot 4. latents converge quicker for "easier" tokens, early looping exit is thus possible per token, depending on how "hard" the token is 5. takes up less memory in training/inference (tnx to truncated BPTT and KV sharing respectively) few mechinterp notes: 1. might be harder to do J-space/J-lens type of analysis (the deeper the model the harder it is to trust the jacobians that map intermediate activations to the output); in general this leads to more "neauralese" and makes it harder to solely use cot to monitor the models (i suspect that's what we've been seeing from oai/ant recently) 2. they find evidence for interesting structures in latent space, such as the tendency of the latent representation to follow "orbitals" to compute arithmetic tasks and for reasoning in general (the shape rotators win again?) -- in general there is less need for long cot data, but it’s fully complementary one can and will do both in practice another example you can do great work at academia! soon we’ll have incredibly strong OSS models that can run locally, we’ve decoupled num of params from the perf of the model! intelligence will be abundant, aligned to you, and democratized :)21d
    Andrew Gordon Wilson@andrewgwilsTypically, it's assumed that only data interventions can affect scaling exponents, while the architecture and optimizer at best can change the constants. And indeed, until now, data has been the primary driver of pre-training efficiency. 2/720d

    8 Sources

    SemiAnalysis@SemiAnalysis_GPT-6 Astra reportedly gets deeper without getting bigger. The labs already know what that means for scaling. "It's basically confirmed that GPT-6 Astra uses loop transformers, which means that instead of adding parameter count, it goes through the layers more than once. You add compute depth, but you don't increase the size of the model." "The labs are probably the people that are best positioned to say which way models are scaling. I think that's a tell that they're not seeing parameter sizes scaling as aggressively in their roadmaps, in what they find in their research."23d
    Aleksa Gordić (水平问题)@gordic_aleksalooped transformer unlocks a new scaling dimension - recurrent depth; i think it received less love (other than the rumors that frontier ai labs use it) than it deserves, here are few notes / observations fundamentally the method lets the model think deeper without increasing either its parameter count or its verbalized chain of thought length (the other 2 important scaling dimensions) i find the constrast with moe interesting: > looped transformer: keep params const, increase total FLOPs (more computation = reasoning) > moe: keep FLOPs const (through active params), increase total params (more param capacity = memory/knowledge) the architectural modification is straight forward, instead of: L1 → L2 → … → L_n (where n = num layers) we do: P → R → R → … → R → C where: * P - prelude, few transformer layers that we apply only once at the beginning (i.e. L1 -> L2 -> ... -> L_p) * R - recurring core, normally this would constitute the middle layers of the transformer, but in constrast with regular transformer R is repeated/looped an arbitrary number of times (r) * C - postlude or "coda", similar to P just at the exit so the computation proceeds as following: 1. e = P(x) 2. apply s_i = R(e, s_i-1) r times, where s0 ~ N(0, sigma^2 * I), and s and e are concat and mapped/"projected" back to hidden size 3. apply C on the final s (state) (the particularly important design decision is that e is re-injected at every recurrence) now the beautiful part is that more r at inference improves perf, and harder / more context-heavy problems benefit from more latent compute. e.g. on ARC-C they demonstrate: * 0-shot -> perf saturates around 8-12 iterations * 1-shot -> 20 iters * few shot (25-50 examples) -> 32 iterations (the recurrent reasoning is like a learned optimization in the latent space during inference; grad-descent during inference unlocked?) -- the above generalization is downstream of how we train the model, r is sampled from heavy tailed log-normal-Poisson distribution during training thus the model is trained across randomized recurrent depths (they use truncated BPTT - backprop through depth in this case - to save on memory) -- few systems notes: 1. they synchronize one r per microbatch across workers, otherwise GPUs with small r would be idle waiting for workers that happened to sample large r 2. the model has a natural specdec capability -> draft model == low recurrence count (r), verifier model == high recurrence count. states computed during drafting can be reused when verifying! e.g. s0→s1→s2→s3 from draft (r=4) can be fed into verifier => s3→s4→…→s11 (r=12) 3. aggresive KV cache sharing is possible because every recurrent iteration uses the same weights, they show that caches from different recurrence depths can be aggressively reused zero-shot 4. latents converge quicker for "easier" tokens, early looping exit is thus possible per token, depending on how "hard" the token is 5. takes up less memory in training/inference (tnx to truncated BPTT and KV sharing respectively) few mechinterp notes: 1. might be harder to do J-space/J-lens type of analysis (the deeper the model the harder it is to trust the jacobians that map intermediate activations to the output); in general this leads to more "neauralese" and makes it harder to solely use cot to monitor the models (i suspect that's what we've been seeing from oai/ant recently) 2. they find evidence for interesting structures in latent space, such as the tendency of the latent representation to follow "orbitals" to compute arithmetic tasks and for reasoning in general (the shape rotators win again?) -- in general there is less need for long cot data, but it’s fully complementary one can and will do both in practice another example you can do great work at academia! soon we’ll have incredibly strong OSS models that can run locally, we’ve decoupled num of params from the perf of the model! intelligence will be abundant, aligned to you, and democratized :)21d
    Andrew Gordon Wilson@andrewgwilsTypically, it's assumed that only data interventions can affect scaling exponents, while the architecture and optimizer at best can change the constants. And indeed, until now, data has been the primary driver of pre-training efficiency. 2/720d