• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Byte-level models reportedly pull ahead as compute grows in Meta research

    A post discussing a Meta paper says its distilled 1-billion-parameter byte models matched a distilled token model with one-sixth of the training data.

    CS
    EL
    ML
    3 Sources, ,

    TLDR

    A post about a Meta paper describes distilled 1-billion-parameter models trained on up to 1 trillion bytes. It says token models lead at low compute but plateau, while byte-level models overtake them as compute grows. According to the post, the work trains byte-level students by converting token-based teachers’ output scores into byte-level scores. Its exact conversion method is called End-Of-Token. The post says fitted scaling laws predict that model will end up to 4% ahead of the distilled token model. It also reports that a 256-entry vocabulary cuts storage for teacher output scores to about a fifth.

    Combined views

    54.6K

    3 Sources, first seen 16d ago

    Combined views

    54.6K

    3 Sources, first seen 16d ago

    855 likes
    16d ago
    first seen 16d ago
    855 likes
    22 comments
    826 saves
    209 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    22 comments
    826 saves
    209 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @omarsar0Banger paper from Meta. This work shows that byte-level models start out behind token models and then pass them as compute grows. They show this for distilled 1B models trained on up to 1 trillion bytes. To distill a byte student from a token teacher, they convert the teacher's token logits into byte logits, either approximately (Marginalize-It) or exactly (End-Of-Token). Token models lead at low compute but plateau. Byte models reach a higher ceiling, and the fitted scaling laws predict the End-Of-Token model ends up to 4% ahead of the distilled token model. The byte models also match the distilled token model with one-sixth of the training data, and a 256-entry vocabulary cuts teacher-logit storage to about a fifth. Paper: https://arxiv.org/abs/2609.12303 Chat with Paper: https://academy.dair.ai/papers/breaking-the-token-ceiling-distilling-smaller-stronger-byte-models-2609.12303
    @ChrSzegedyRT @omarsar0: Banger paper from Meta. This work shows that byte-level models start out behind token models and then pass them as compute g…
    @MLStreetTalkRT @omarsar0: Banger paper from Meta. This work shows that byte-level models start out behind token models and then pass them as compute g…

    3 Sources

    @omarsar0Banger paper from Meta. This work shows that byte-level models start out behind token models and then pass them as compute grows. They show this for distilled 1B models trained on up to 1 trillion bytes. To distill a byte student from a token teacher, they convert the teacher's token logits into byte logits, either approximately (Marginalize-It) or exactly (End-Of-Token). Token models lead at low compute but plateau. Byte models reach a higher ceiling, and the fitted scaling laws predict the End-Of-Token model ends up to 4% ahead of the distilled token model. The byte models also match the distilled token model with one-sixth of the training data, and a 256-entry vocabulary cuts teacher-logit storage to about a fifth. Paper: https://arxiv.org/abs/2609.12303 Chat with Paper: https://academy.dair.ai/papers/breaking-the-token-ceiling-distilling-smaller-stronger-byte-models-2609.12303
    @ChrSzegedyRT @omarsar0: Banger paper from Meta. This work shows that byte-level models start out behind token models and then pass them as compute g…
    @MLStreetTalkRT @omarsar0: Banger paper from Meta. This work shows that byte-level models start out behind token models and then pass them as compute g…