• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Elie Bakouch Questions Mistral GLM 5.3 Serving Time

    Research engineer notes GLM 5.3 shares architecture with 5.2.

    AK
    MA
    EL
    10 Sources, 30d ago, first seen 30d ago

    TLDR

    Elie Bakouch, an ML research engineer, posted that Mistral serving GLM 5.3 depends on whether the company makes modifications. He stated the model uses the exact same architecture as GLM 5.2. Bakouch added that without changes it amounts to swapping weights, described as a one-line code change. He mentioned possible extra steps such as nvfp4 quantization or training dspark or flash kernels. The post appears in a cluster tagged Research and includes a generated headline reference to GLM 5.2 in Europe.

    Combined views

    75K

    10 Sources, first seen 30d ago

    Combined views

    75K

    10 Sources, first seen 30d ago

    572 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    572 likes
    55 comments
    99 saves
    35 reposts
    55 comments
    99 saves
    35 reposts

    10 Sources

    @eliebakouchvery nice but now the real question is: how long is it going to take for mistral to serve glm 5.3? this is the exact same architecture as 5.2, so if they don't make any modifications, it's just swapping the weights which is a 1 line code change if they do some additional nvfp4 quantization or train some dspark/flash, might take a few days in theory to apply the same pipeline (glm 5.3 has been released for 5 days) also there is no public information about tok/s and latency that mistral hosting offers? mistral docs is empty + they mention GLM 5.2 at the same level of "performance" as mistral medium 3.5 or large 3 which i imagine is super misleading if you don't know much about ai models (they say "third party open source model" for glm5.2 and "frontier-class" and "state-of-the-art" for mistral models lol) anyway, positive news overall since corporate people that are forced to use mistral just got a massive upgrade, but as always they are lagging behind in a way
    @xlr8harderGiven their existing work in in compression/quantization techniques and this changelog entry, is this just glm 5.2? https://docs.compactif.ai/changelog/
    @American_Gladio@casper_hansen_ It’s just GLM 5.2
    @zehavocRT @American_Gladio: @casper_hansen_ It’s just GLM 5.2
    @BlackHCCompressed GLM 5.2 tho tbh H/t @xlr8harder
    @DorialexanderSo not Nemotron. Between this and Mistral, I guess it’s good sovereign EU is rediscovering GLM 5.2 in one shape of another three months later.
    @artetxemOnce upon a time, a Spanish startup launched the "Zetta Multiverso", supposedly the first smartphone from Extremadura (a region of Spain). Turns out it was a Xiaomi phone with their sticker on it. Chinese technology marketed as Spanish. Now Multiverse is launching Quasar, "the top European AI model". Turns out it's compressed GLM-5.2. Chinese technology marketed as European. Zetta Multiverso. Multiverse Quasar. Maybe it's something about the name.

    Sentiment

    Positive3.3%96.7%Negative

    Summary

    Sentiment

    Positive3.3%96.7%Negative

    Replies dismissed Mistral models as useless toys made by incompetents and mocked GLM 5.2 as a rebranded compressed model, while criticizing CNRS for forcing its use despite researchers preferring other services.

    Based on 44 sentiment-bearing replies from 30 accounts across 2 conversations.

    Summary

    Replies dismissed Mistral models as useless toys made by incompetents and mocked GLM 5.2 as a rebranded compressed model, while criticizing CNRS for forcing its use despite researchers preferring other services.

    Based on 44 sentiment-bearing replies from 30 accounts across 2 conversations.

    10 Sources

    @eliebakouchvery nice but now the real question is: how long is it going to take for mistral to serve glm 5.3? this is the exact same architecture as 5.2, so if they don't make any modifications, it's just swapping the weights which is a 1 line code change if they do some additional nvfp4 quantization or train some dspark/flash, might take a few days in theory to apply the same pipeline (glm 5.3 has been released for 5 days) also there is no public information about tok/s and latency that mistral hosting offers? mistral docs is empty + they mention GLM 5.2 at the same level of "performance" as mistral medium 3.5 or large 3 which i imagine is super misleading if you don't know much about ai models (they say "third party open source model" for glm5.2 and "frontier-class" and "state-of-the-art" for mistral models lol) anyway, positive news overall since corporate people that are forced to use mistral just got a massive upgrade, but as always they are lagging behind in a way
    @xlr8harderGiven their existing work in in compression/quantization techniques and this changelog entry, is this just glm 5.2? https://docs.compactif.ai/changelog/
    @American_Gladio@casper_hansen_ It’s just GLM 5.2
    @zehavocRT @American_Gladio: @casper_hansen_ It’s just GLM 5.2
    @BlackHCCompressed GLM 5.2 tho tbh H/t @xlr8harder
    @DorialexanderSo not Nemotron. Between this and Mistral, I guess it’s good sovereign EU is rediscovering GLM 5.2 in one shape of another three months later.
    @artetxemOnce upon a time, a Spanish startup launched the "Zetta Multiverso", supposedly the first smartphone from Extremadura (a region of Spain). Turns out it was a Xiaomi phone with their sticker on it. Chinese technology marketed as Spanish. Now Multiverse is launching Quasar, "the top European AI model". Turns out it's compressed GLM-5.2. Chinese technology marketed as European. Zetta Multiverso. Multiverse Quasar. Maybe it's something about the name.