• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    User reports halving Qwen3-8B’s prefill time with similar output

    The user says an added approximation model sped up input processing without changing Qwen3-8B itself, using a technique they attribute to DeepSeek-V4.1-Flash.

    T(
    きし
    2 Sources, 20d ago, first seen 20d ago

    TLDR

    A user reports applying a technique they say DeepSeek-V4.1-Flash calls “Encoder-Decoder” to Qwen3-8B, approximating key-value (KV) data in the model’s later layers. They say this halved prefill time—the time spent processing input before generating output—with similar output. Their approach adds an approximation model rather than changing Qwen3-8B itself, according to the user.

    Combined views

    201.1K

    2 Sources, first seen 20d ago

    Combined views

    201.1K

    2 Sources, first seen 20d ago

    1.2K likes
    1.2K likes
    25 comments
    752 saves
    190 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    25 comments
    752 saves
    190 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @kisやった! DeepSeek-V4.1-FlashがEncoder-Decoderと呼んでいる後半層KV近似の仕組みをQwen3-8Bに適用して、プレフィル時間半分、出力同様というのができた! つまりこの仕組みは、既存のモデルでも、モデル自体をいじらず近似モデルを追加してあげれば、プレフィル時間を半分にすることができる。
    @teortaxesTexYou can just use normally trained decodes in the YOCO regime?

    2 Sources

    @kisやった! DeepSeek-V4.1-FlashがEncoder-Decoderと呼んでいる後半層KV近似の仕組みをQwen3-8Bに適用して、プレフィル時間半分、出力同様というのができた! つまりこの仕組みは、既存のモデルでも、モデル自体をいじらず近似モデルを追加してあげれば、プレフィル時間を半分にすることができる。
    @teortaxesTexYou can just use normally trained decodes in the YOCO regime?