• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    E2S finetuning's claimed gains over two training baselines on new tasks

    Its creator says E2S uses GFlowNets to turn expert demonstrations into training data better aligned with a student model.

    Lianhui Qin@COLM2026LQ
    Murray KangMK
    2 Sources, ,

    TLDR

    The developer describes E2S as a way to generate supervised fine-tuning examples that preserve information from expert demonstrations while better matching how a student model responds. They report gains of 6.0 points over MCMC+SFT and 35.2 points over vanilla expert SFT on new tasks, while preserving performance on prior tasks.

    Combined views

    950

    2 Sources, first seen 5h ago

    Combined views

    950

    2 Sources, first seen 5h ago

    25 likes
    5h ago
    first seen 5h ago
    25 likes
    1 comments
    15 saves
    4 reposts
    1 comments
    15 saves
    4 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    Murray Kang@haoqik322Make SFT great again! ✊ Excited to share my weekend project: E2S (Expert-to-Student) Finetuning — an amortized sampling approach for turning off-policy expert demonstrations into more on-policy SFT data. The problem: vanilla SFT asks the student model to imitate correct but off-policy expert trajectories, causing unnecessary policy shifts and forgetting. Our idea is to train a model to speak the expert data in the student’s own language. We formulate this as a constrained generation problem: the expert defines the constraint — what privileged information should be preserved — while the student decides how to generate a more "on-policy" response under that constraint. The recent Finetuning with Sampling paper targets this distribution with per-example MCMC. But inference-time searching for every example is expensive. E2S amortizes the search: learn a conditional sampler across examples with GFlowNets, then reuse it to generate more on-policy SFT targets. On new tasks, E2S gains +6.0 points over MCMC+SFT and +35.2 points over vanilla expert SFT, while preserving performance on prior tasks. See details in the blog: https://mk322.github.io/blog/learned-on-policy-sampler/5h
    Lianhui Qin@COLM2026@LianhuiqRT @haoqik322: Make SFT great again! ✊ Excited to share my weekend project: E2S (Expert-to-Student) Finetuning — an amortized sampling app…2h

    2 Sources

    Murray Kang@haoqik322Make SFT great again! ✊ Excited to share my weekend project: E2S (Expert-to-Student) Finetuning — an amortized sampling approach for turning off-policy expert demonstrations into more on-policy SFT data. The problem: vanilla SFT asks the student model to imitate correct but off-policy expert trajectories, causing unnecessary policy shifts and forgetting. Our idea is to train a model to speak the expert data in the student’s own language. We formulate this as a constrained generation problem: the expert defines the constraint — what privileged information should be preserved — while the student decides how to generate a more "on-policy" response under that constraint. The recent Finetuning with Sampling paper targets this distribution with per-example MCMC. But inference-time searching for every example is expensive. E2S amortizes the search: learn a conditional sampler across examples with GFlowNets, then reuse it to generate more on-policy SFT targets. On new tasks, E2S gains +6.0 points over MCMC+SFT and +35.2 points over vanilla expert SFT, while preserving performance on prior tasks. See details in the blog: https://mk322.github.io/blog/learned-on-policy-sampler/5h
    Lianhui Qin@COLM2026@LianhuiqRT @haoqik322: Make SFT great again! ✊ Excited to share my weekend project: E2S (Expert-to-Student) Finetuning — an amortized sampling app…2h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet