• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Elie Bakouch Speculates on Model Distillation

    ML engineer Elie Bakouch replies to @scaling01 expressing uncertainty about a model.

    EL
    LA
    6 Sources, 29d ago, first seen 29d ago

    TLDR

    Research engineer Elie Bakouch, who trains LLMs at Prime Intellect after contributing to SmolLM and FineWeb at Hugging Face, posted a reply to user @scaling01. He wrote that he now sees how the point makes more sense yet remains unsure about the matter. Bakouch added that he would think the item is distilled from model 1 or 2 likely. The comment carries a research tag and forms part of visible replies in the conversation.

    Combined views

    5.4K

    6 Sources, first seen 29d ago

    Combined views

    5.4K

    6 Sources, first seen 29d ago

    50 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    50 likes
    7 comments
    2 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 comments
    2 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 Sources

    @scaling01@eliebakouch that is not supported: "Fable is not distilled version of a larger Mythos model" we already knew before, that the public Fable and Mythos are the same size and arch the claim is that they have a larger "Mythos Preview" kind of teacher from which Mythos/Fable is derived
    @eliebakouch@scaling01 i think i've heard that for mythos 5 -> fable 5, also what the reasoning behind mythos preview as bigger model? do we have benchmark where it perform much better?

    6 Sources

    @scaling01@eliebakouch that is not supported: "Fable is not distilled version of a larger Mythos model" we already knew before, that the public Fable and Mythos are the same size and arch the claim is that they have a larger "Mythos Preview" kind of teacher from which Mythos/Fable is derived
    @eliebakouch@scaling01 i think i've heard that for mythos 5 -> fable 5, also what the reasoning behind mythos preview as bigger model? do we have benchmark where it perform much better?