• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Masked Distillation is claimed to eliminate LLM thinking-token latency on particular distributions

    A post questions whether cutting the average length of an LLM’s intermediate-token output reflects greater efficiency or overtraining of the base model.

    Subbarao Kambhampati (కంభంపాటి సుబ్బారావు)SK
    2 Sources, 21d ago, first seen 21d ago

    TLDR

    A September 16 post claims that over-training with Masked Distillation can eliminate delays caused by generating intermediate “thinking” tokens on particular distributions. The author presents it as a technique from their own work. In a July 30 post, the same author questioned research that equates shorter average intermediate-token output with improved efficiency, and described work investigating those claims through Masked Distillation.

    Combined views

    3.4K

    2 Sources, first seen 21d ago

    Combined views

    3.4K

    2 Sources, first seen 21d ago

    17 likes
    17 likes
    2 comments
    19 saves
    6 reposts
    2 comments
    19 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Subbarao Kambhampati (కంభంపాటి సుబ్బారావు)@rao2zYou can get rid of think token latency by over-training with Masked Distillation (aka you can always compile System 2 to System 1): Read the LLM du jour press release--the "System 1 LLM" named after Jevon with interest. While it is hard to to figure out the technical ideas from press releases, if all you wanted was to get rid of latency induced by think tokens on particular distributions, you can just use our Masked Distillation technique. 👇21d

    2 Sources

    Subbarao Kambhampati (కంభంపాటి సుబ్బారావు)@rao2zYou can get rid of think token latency by over-training with Masked Distillation (aka you can always compile System 2 to System 1): Read the LLM du jour press release--the "System 1 LLM" named after Jevon with interest. While it is hard to to figure out the technical ideas from press releases, if all you wanted was to get rid of latency induced by think tokens on particular distributions, you can just use our Masked Distillation technique. 👇21d