• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Jie Tang on Scaling Laws Beyond Parameters

    Professor highlights role of data, compute allocation, and inference in model performance.

    ND
    TW
    LF
    13 Sources, 43d ago, first seen 43d ago

    TLDR

    Jie Tang, Tsinghua professor and Zhipu AI co-founder, posted that parameter count alone does not answer how capable a model will be. He stated it must be weighed with data volume, compute allocation choices, and inference demands. The post drew replies from Nando de Freitas, Thomas Wolf, Julian Togelius, Delip Rao, and others who described it as a clear account of current scaling trade-offs and a masterclass on the subject.

    Combined views

    1.3M

    13 Sources, first seen 43d ago

    Combined views

    1.3M

    13 Sources, first seen 43d ago

    6.8K likes
    6.8K likes
    238 comments
    4.4K saves
    756 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    238 comments
    4.4K saves
    756 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    13 Sources

    @jietangThoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
    @teortaxesTex«Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim»
    @kimmonismusInteresting take by the zAI (GLM-Models) founder: AI scaling is not over, but we just focused too much on model size. For years, the industry mostly asked: How many parameters does a model have? But performance also depends on training data, inference compute and post-training. GLM-5.3 is a very good example here. It reportedly uses the same base model, architecture and parameter count as GLM-5.2. The main change was one month of additional reinforcement learning, and the gains were very substantial. In short: the focus should be placed more on the other potentials and not just on the size.
    @delipraoThis is a masterclass in scaling laws, and also on how to share research outcomes.
    @Thom_WolfI wish every neolab had a professor as cofounder of the level of @jietang and so able to put in perspective their new model release. A great snapshot on the history of scaling laws
    @samsja19assuming we can scale active and total independently. What is the target in active parameters (and depth to some extend) you feel are enough to reach the next stage of capabilities ? Also in term of depth, what do you think about the trade off of scaling depth / active vs test time compute ? is there any reason not to scale total parameters really high tho ? like should we always max out the experts count assuming it fit well in memory / ep is still fast on nv72 ?
    @togeliusI learned more from reading this short post from the founder of Z than from all the posts I’ve read from Amodei and Altman
    @nxthompsonThis analysis of the different ways to improve models — and to optimize for solving different bottlenecks at different times — is so smart and so clear. Anyone who says the Chinese labs are catching up solely by copying the American labs is making a mistake.
    @NandoDFThis is the wisest and most accurate take on scaling laws — their promise and their costly failures — that you could read today. This is what building AI systems is currently all about. Well said @jietang
    @LiamFedusAn excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!

    13 Sources

    @jietangThoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
    @teortaxesTex«Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim»
    @kimmonismusInteresting take by the zAI (GLM-Models) founder: AI scaling is not over, but we just focused too much on model size. For years, the industry mostly asked: How many parameters does a model have? But performance also depends on training data, inference compute and post-training. GLM-5.3 is a very good example here. It reportedly uses the same base model, architecture and parameter count as GLM-5.2. The main change was one month of additional reinforcement learning, and the gains were very substantial. In short: the focus should be placed more on the other potentials and not just on the size.
    @delipraoThis is a masterclass in scaling laws, and also on how to share research outcomes.
    @Thom_WolfI wish every neolab had a professor as cofounder of the level of @jietang and so able to put in perspective their new model release. A great snapshot on the history of scaling laws
    @samsja19assuming we can scale active and total independently. What is the target in active parameters (and depth to some extend) you feel are enough to reach the next stage of capabilities ? Also in term of depth, what do you think about the trade off of scaling depth / active vs test time compute ? is there any reason not to scale total parameters really high tho ? like should we always max out the experts count assuming it fit well in memory / ep is still fast on nv72 ?
    @togeliusI learned more from reading this short post from the founder of Z than from all the posts I’ve read from Amodei and Altman
    @nxthompsonThis analysis of the different ways to improve models — and to optimize for solving different bottlenecks at different times — is so smart and so clear. Anyone who says the Chinese labs are catching up solely by copying the American labs is making a mistake.
    @NandoDFThis is the wisest and most accurate take on scaling laws — their promise and their costly failures — that you could read today. This is what building AI systems is currently all about. Well said @jietang
    @LiamFedusAn excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!