• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    How multi-teacher on-policy distillation combines specialist capabilities in one model

    A post explains that the student model generates text, then a teacher for that domain scores it token by token.

    Cameron R. Wolfe, Ph.D.CR
    1 Source, ,

    TLDR

    A post describes multi-teacher on-policy distillation (MOPD) as a way to train specialist models separately, then consolidate their capabilities into one student model. The student generates text, and a teacher for the relevant domain scores each token rather than generating the text itself. The post says this dense feedback can make distillation more sample-efficient than outcome-based reinforcement learning, but cautions that teachers whose behavior differs too much from the student's may provide a weaker learning signal.

    Combined views

    2.1K

    1 Source, first seen 1h ago

    Combined views

    2.1K

    1 Source, first seen 1h ago

    42 likes
    1h ago
    first seen 1h ago
    42 likes
    8 comments
    42 saves
    8 reposts
    8 comments
    42 saves
    8 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Cameron R. Wolfe, Ph.D.@cwolferesearchHere how multi-teacher on-policy distillation (MOPD) works… TL;DR: Multi-Teacher On-Policy Distillation (MOPD) provides a way to consolidate capabilities from several specialized teacher models into a single student. Rather than training one policy to become an expert in many domains simultaneously, we can train specialists independently (often in parallel) and distill them back into one model using dense, token-level supervision. On-policy distillation (OPD) basics. MOPD builds upon the idea of OPD. Importantly, OPD is on-policy from the perspective of the student. Given a prompt, the student (not the teacher!) generates a rollout from this prompt. Then, the teacher computes probabilities for each student-generated token in this rollout, and we train the student to match the teacher using a (sampled) reverse-KL objective. Put simply, we use the teacher to score the observed behavior of the student, rather than actually generating a trajectory with the teacher. Extension to MOPD. MOPD extends this setup to multiple teachers. Instead of scoring every rollout with the same teacher, each prompt / rollout is routed to the teacher that specializes in the corresponding domain for that prompt. A SWE rollout will be scored by a SWE teacher, a chat prompt is routed to a conversational teacher, and so on. Dense supervision. One major benefit of OPD is that it provides a dense learning signal. Outcome-based RL (most common RL training setting for LLMs) assigns one reward to a full trajectory, whereas OPD provides supervision at every generated token. This dense signal can make distillation more sample efficient relative to outcome-based RL. However, OPD / MOPD and RL are not mutually exclusive—we can combine the distillation objective with a standard policy-gradient objective and learn from both teacher probabilities and task-specific rewards. Specialized teachers. The teachers used for MOPD can be independently optimized for different capabilities (e.g., reasoning, coding, SWE, search, etc.). Teachers can use their own post-training pipeline—combining SFT, RLVR, RLHF, and more depending on what works best for the domain. For example, Nemotron 3 Ultra trains ~10 teacher models by applying bespoke training pipelines (with different combinations of post-training algorithms); see image. Even though teachers often use their own training pipelines, teachers cannot be combined arbitrarily: if a teacher's behavior differs too much from the student, student-generated rollouts may be out-of-distribution for that teacher, weakening the supervision signal. This can be solved via things like MOPD warmup, where we train the student with SFT on teacher rollouts prior to MOPD to get student behavior closer to that of the teachers. Why is this helpful? Scaling RL across many domains is difficult. As we add domains, each capability receives fewer rollouts, domains can interfere with one another, and differences in rollout length / verification cost make training infrastructure increasingly complicated. MOPD gives us another approach: train domain specialists independently (often in parallel) then consolidate their capabilities into a single policy. It can also be used to recover capabilities that regress during sequential stages of post-training.1h

    1 Source

    Cameron R. Wolfe, Ph.D.@cwolferesearchHere how multi-teacher on-policy distillation (MOPD) works… TL;DR: Multi-Teacher On-Policy Distillation (MOPD) provides a way to consolidate capabilities from several specialized teacher models into a single student. Rather than training one policy to become an expert in many domains simultaneously, we can train specialists independently (often in parallel) and distill them back into one model using dense, token-level supervision. On-policy distillation (OPD) basics. MOPD builds upon the idea of OPD. Importantly, OPD is on-policy from the perspective of the student. Given a prompt, the student (not the teacher!) generates a rollout from this prompt. Then, the teacher computes probabilities for each student-generated token in this rollout, and we train the student to match the teacher using a (sampled) reverse-KL objective. Put simply, we use the teacher to score the observed behavior of the student, rather than actually generating a trajectory with the teacher. Extension to MOPD. MOPD extends this setup to multiple teachers. Instead of scoring every rollout with the same teacher, each prompt / rollout is routed to the teacher that specializes in the corresponding domain for that prompt. A SWE rollout will be scored by a SWE teacher, a chat prompt is routed to a conversational teacher, and so on. Dense supervision. One major benefit of OPD is that it provides a dense learning signal. Outcome-based RL (most common RL training setting for LLMs) assigns one reward to a full trajectory, whereas OPD provides supervision at every generated token. This dense signal can make distillation more sample efficient relative to outcome-based RL. However, OPD / MOPD and RL are not mutually exclusive—we can combine the distillation objective with a standard policy-gradient objective and learn from both teacher probabilities and task-specific rewards. Specialized teachers. The teachers used for MOPD can be independently optimized for different capabilities (e.g., reasoning, coding, SWE, search, etc.). Teachers can use their own post-training pipeline—combining SFT, RLVR, RLHF, and more depending on what works best for the domain. For example, Nemotron 3 Ultra trains ~10 teacher models by applying bespoke training pipelines (with different combinations of post-training algorithms); see image. Even though teachers often use their own training pipelines, teachers cannot be combined arbitrarily: if a teacher's behavior differs too much from the student, student-generated rollouts may be out-of-distribution for that teacher, weakening the supervision signal. This can be solved via things like MOPD warmup, where we train the student with SFT on teacher rollouts prior to MOPD to get student behavior closer to that of the teachers. Why is this helpful? Scaling RL across many domains is difficult. As we add domains, each capability receives fewer rollouts, domains can interfere with one another, and differences in rollout length / verification cost make training infrastructure increasingly complicated. MOPD gives us another approach: train domain specialists independently (often in parallel) then consolidate their capabilities into a single policy. It can also be used to recover capabilities that regress during sequential stages of post-training.1h