• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GLM-5.3 reportedly helped triple inference throughput for GLM-5.3-Flash

    Z.ai says the system reached production readiness less than two weeks after its first successful run. It credits detailed tests and measurements with guiding the optimization.

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)T(
    Andrew CurranAC
    Rohan PaulRP
    11 Sources, ,

    TLDR

    Z.ai says GLM-5.3 helped build and optimize the infrastructure running GLM-5.3-Flash, with end-to-end throughput tripling relative to the initial baseline. The company reports that the system went from its first successful run to production readiness in less than two weeks. It credits correctness tests, execution traces, microbenchmarks and end-to-end measurements with enabling targeted hypothesis testing rather than reliance on aggregate performance metrics alone.

    Combined views

    1.4M

    11 Sources, first seen 21d ago

    Combined views

    1.4M

    11 Sources, first seen 21d ago

    8.1K likes
    21d ago
    first seen 21d ago
    8.1K likes
    376 comments
    2.9K saves
    804 reposts
    376 comments
    2.9K saves
    804 reposts

    Sentiment

    Positive79.4%20.6%Negative

    Based on 35 sentiment-bearing replies from 34 accounts across 2 conversations.

    Sentiment

    Positive79.4%20.6%Negative

    Based on 35 sentiment-bearing replies from 34 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    11 Sources

    Z.ai@Zai_orgWe’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure21d
    jietang@jietangTwo weeks. That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way. What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3. A model helping optimize the system that serves it. The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode). But the most important lesson wasn't about any single optimization. When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know *why* things got worse. "Throughput down 20%" tells you something broke. It doesn't tell you which layer, which hypothesis, or what to test next. In RL terms, it's a sparse reward with a credit assignment problem. And an end-to-end benchmark that takes hours makes exploration painfully slow. Senior engineers solve this with an implicit process reward in their heads. They know when to check the timeline, when to run a microbenchmark, and which layer's output to compare. So we made that explicit. We call it dense feedback: layered verification interfaces the agent can call directly. Correctness feedback: did it compute right? System behavior feedback: where did the time go? Performance feedback: which option wins, under which conditions? Each signal has to be local, cheap, and objectively verifiable. Three things the agent found: First, precision drift in KDA's context-parallel path that grew with sequence length. The cause was TF32 rounding error compounding through chained state-matrix merges. The fix is now merged upstream in Flash Linear Attention (PR #1180). Second, KV transfer never overlapped with DeepEP dispatch. The agent followed the call chain across the Python/C++ boundary and found that the intranode path never released the GIL. After the fix, transfer overhead fell from over 30% to under 1%. Third, a decode kernel recomputing the same normalization four times because of how it was chunked. The agent restructured it and got a 1.71× speedup. The idea came from "optimization skeletons" it had distilled by reading existing kernels across SGLang, FLA, and DeepGEMM. To be clear about the boundaries: humans still defined the goals, built the feedback environment, and reviewed every high-risk change. But the engineer's role is changing, from the person who solves the problem to the person who designs the feedback. There's a deeper implication too. A layered, verifiable feedback environment built on real infrastructure tasks is exactly what training the next generation of models needs most. Every task the agent completes can become training ground for its successor. We are still far from recursive self-improvement. But the smallest loop now exists. The model optimizes the system. The system serves the model.21d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexThis is not really about RSI, vut this is TCD; this is how the CUDA moat dies. First, in inference optimization. Wenfeng talks about this. Agents are too good now. Every scrap of silicon will be pushed to its limits.20d
    Chubby♨️@kimmonismusWhile the US is discussing a slowdown, China is currently putting full force into its RSI. Zai says GLM is taking early steps toward recursive self-improvement: AI helping build and optimize the infrastructure that runs AI. GLM-5.3 helped build the infrastructure running GLM-5.3-Flash, with engineers and the agent tripling inference throughput *in under two weeks*. They say this work will “directly change how the next generation of models is trained.”20d
    Andrew Curran@AndrewCurran_From Z․ai's research blog post this morning documenting the early signs they are seeing of recursive self improvement. 'GLM is increasingly helping build AI itself'20d
    Rohan Paul@rohanpaul_aiA brilliant post from the GLM-5.3 team on RSI (recursive self-improvement). GLM-5.3 was already used to optimize the infrastructure that runs GLM itself, including production kernel and concurrency fixes. It helped triple GLM-5.3-Flash throughput on 100,000+ accelerators. Their early self-improvement loop: the model improves its serving system, that system runs the model, and the engineering knowledge accumulates for the next optimization cycle. Engineers still set objectives and boundaries, while the agent handled analysis, hypotheses, code changes, and experiments. In one test, Prefill plus KV Transfer lagged Prefill alone by over 20%, and the agent traced the slowdown to the Python GIL, and releasing that lock cut the gap below 1% Another kernel change reached a 1.71x speedup over the prior version by eliminating repeated FP32 normalization and gating work.20d
    Minh Nhat Nguyen 🦭@menhguinZAI and Minimax will be the first publicly traded companies where recursive-self-improvement will materially affect earnings quarter-on-quarter.20d
    Josh Constine 📶🔥@JoshConstine@0xSigil Wait, you're somehow running an inference optimization lab inside a consumer agent startup? Team must be cracked16d

    11 Sources

    Z.ai@Zai_orgWe’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure21d
    jietang@jietangTwo weeks. That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way. What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3. A model helping optimize the system that serves it. The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode). But the most important lesson wasn't about any single optimization. When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know *why* things got worse. "Throughput down 20%" tells you something broke. It doesn't tell you which layer, which hypothesis, or what to test next. In RL terms, it's a sparse reward with a credit assignment problem. And an end-to-end benchmark that takes hours makes exploration painfully slow. Senior engineers solve this with an implicit process reward in their heads. They know when to check the timeline, when to run a microbenchmark, and which layer's output to compare. So we made that explicit. We call it dense feedback: layered verification interfaces the agent can call directly. Correctness feedback: did it compute right? System behavior feedback: where did the time go? Performance feedback: which option wins, under which conditions? Each signal has to be local, cheap, and objectively verifiable. Three things the agent found: First, precision drift in KDA's context-parallel path that grew with sequence length. The cause was TF32 rounding error compounding through chained state-matrix merges. The fix is now merged upstream in Flash Linear Attention (PR #1180). Second, KV transfer never overlapped with DeepEP dispatch. The agent followed the call chain across the Python/C++ boundary and found that the intranode path never released the GIL. After the fix, transfer overhead fell from over 30% to under 1%. Third, a decode kernel recomputing the same normalization four times because of how it was chunked. The agent restructured it and got a 1.71× speedup. The idea came from "optimization skeletons" it had distilled by reading existing kernels across SGLang, FLA, and DeepGEMM. To be clear about the boundaries: humans still defined the goals, built the feedback environment, and reviewed every high-risk change. But the engineer's role is changing, from the person who solves the problem to the person who designs the feedback. There's a deeper implication too. A layered, verifiable feedback environment built on real infrastructure tasks is exactly what training the next generation of models needs most. Every task the agent completes can become training ground for its successor. We are still far from recursive self-improvement. But the smallest loop now exists. The model optimizes the system. The system serves the model.21d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexThis is not really about RSI, vut this is TCD; this is how the CUDA moat dies. First, in inference optimization. Wenfeng talks about this. Agents are too good now. Every scrap of silicon will be pushed to its limits.20d
    Chubby♨️@kimmonismusWhile the US is discussing a slowdown, China is currently putting full force into its RSI. Zai says GLM is taking early steps toward recursive self-improvement: AI helping build and optimize the infrastructure that runs AI. GLM-5.3 helped build the infrastructure running GLM-5.3-Flash, with engineers and the agent tripling inference throughput *in under two weeks*. They say this work will “directly change how the next generation of models is trained.”20d
    Andrew Curran@AndrewCurran_From Z․ai's research blog post this morning documenting the early signs they are seeing of recursive self improvement. 'GLM is increasingly helping build AI itself'20d
    Rohan Paul@rohanpaul_aiA brilliant post from the GLM-5.3 team on RSI (recursive self-improvement). GLM-5.3 was already used to optimize the infrastructure that runs GLM itself, including production kernel and concurrency fixes. It helped triple GLM-5.3-Flash throughput on 100,000+ accelerators. Their early self-improvement loop: the model improves its serving system, that system runs the model, and the engineering knowledge accumulates for the next optimization cycle. Engineers still set objectives and boundaries, while the agent handled analysis, hypotheses, code changes, and experiments. In one test, Prefill plus KV Transfer lagged Prefill alone by over 20%, and the agent traced the slowdown to the Python GIL, and releasing that lock cut the gap below 1% Another kernel change reached a 1.71x speedup over the prior version by eliminating repeated FP32 normalization and gating work.20d
    Minh Nhat Nguyen 🦭@menhguinZAI and Minimax will be the first publicly traded companies where recursive-self-improvement will materially affect earnings quarter-on-quarter.20d
    Josh Constine 📶🔥@JoshConstine@0xSigil Wait, you're somehow running an inference optimization lab inside a consumer agent startup? Team must be cracked16d