• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Zephyr Argues AI Models Keep Growing in Parameters

    Replies that parameter counts have trended upward with optimization, citing GPT-4 details.

    ZE
    1 Source, 29d ago, first seen 29d ago

    TLDR

    Zephyr states there has never been a trend toward fewer parameters in AI models. The trend has always been toward optimization plus more parameters. The post claims GPT-4 released in 2023 was 1.8T total with 220B active parameters. It adds that model sizes then declined progressively until early 2026. The first production-grade model to cross 2T is mentioned without further detail in the visible text. All points come from this single reply by the pseudonymous commentator.

    Combined views

    27.1K

    1 Source, first seen 29d ago

    Combined views

    27.1K

    1 Source, first seen 29d ago

    98 likes
    98 likes
    9 comments
    33 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    33 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @zephyr_z9"There has never been a trend toward fewer parameters. The trend has always been toward optimization + more parameters" I don't know if u know this, but GPT-4, released in 2023, was 1.8T 220B active and then we progressively saw a decline in model sizes till early 2026 The first production-grade model to cross 2T was Fable/Spud, and now we have Astra, Kimi, and Qwen 3.8 Max (2T+) "Anyway, I really don’t see a world where models stop scaling" I'm not arguing this My argument is about the magnitude of scaling required Will we scale to 100T parameters (which gets deployed widely), or is a 25T looped transformer enough As for KV cache, the size of the KV cache per token has been going down by 10x (at least for Deepseek) YoY, but the context length offered is steadily increasing from 32k to 100k to 256k to 1M and so on A similar trend doesn't exist for parameter scaling

    1 Source

    @zephyr_z9"There has never been a trend toward fewer parameters. The trend has always been toward optimization + more parameters" I don't know if u know this, but GPT-4, released in 2023, was 1.8T 220B active and then we progressively saw a decline in model sizes till early 2026 The first production-grade model to cross 2T was Fable/Spud, and now we have Astra, Kimi, and Qwen 3.8 Max (2T+) "Anyway, I really don’t see a world where models stop scaling" I'm not arguing this My argument is about the magnitude of scaling required Will we scale to 100T parameters (which gets deployed widely), or is a 25T looped transformer enough As for KV cache, the size of the KV cache per token has been going down by 10x (at least for Deepseek) YoY, but the context length offered is steadily increasing from 32k to 100k to 256k to 1M and so on A similar trend doesn't exist for parameter scaling