• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Straightforward multi-teacher distillation may falter when teacher training pipelines differ substantially

    A post discussing the Nemotron 3 Ultra tech report suggests a brief supervised fine-tuning warmup to better align the student with its teachers.

    T(
    CR
    2 Sources, ,

    TLDR

    A post quotes a finding from the Nemotron 3 Ultra tech report: in its trials, teachers trained through substantially different pipelines could not be effectively combined in a straightforward multi-teacher on-policy distillation (MOPD) merge. The post explains that MOPD trains a student model using its own outputs and feedback from teacher models. It suggests first fine-tuning the student on examples from the teachers to make that feedback more useful.

    Combined views

    8.5K

    2 Sources, first seen 7h ago

    Combined views

    8.5K

    2 Sources, first seen 7h ago

    101 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    7h ago
    first seen 7h ago
    101 likes
    7 comments
    77 saves
    8 reposts
    Featured Source
    7 comments
    77 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @cwolferesearchOne of the most interesting details I learned from the Nemotron tech reports was that certain teachers are not compatible to use in MOPD together... Multi-Teacher On-Policy Distillation (MOPD) is a distillation technique that trains a student model on its own rollouts. Given a prompt, we: - Generate a completion from the student. - Route the prompt + completion to an expert / teacher model. - Compute token-level log probabilities from the teacher. Then, we can train the student model using a reverse-KL objective (usually a sampled approximation) between student and teacher probabilities to distill the capabilities of the teacher(s) back into the student policy. Why is this useful? There are a few different setups that can be used for MOPD that differ in how they define / create the teacher models: (1) Consolidation: if we have domain-specific specialist models that have each been created using their own training pipeline, we can consolidate their advanced capabilities into the student with MOPD. This is a great way to improve the capabilities across a broad range of domains, and specialist training is independent across domains (as opposed to multi-domain RL). (2) Recovery: if we are running a sequential / multi-stage training process, we can use checkpoints from earlier training stages as teachers to restore capabilities that regress as we continue to progress through the training process. For either of these cases, we learn from Nemotron 3 Ultra that we cannot just use an arbitrary set of models as teachers for MOPD. See the following quote: “One key finding from our MOPD trials is that teacher models trained with substantially different training pipelines cannot be effectively combined through a straightforward MOPD merge, resulting in suboptimal performance.” - Nemotron 3 Ultra Tech Report MOPD teachers need to be sufficiently compatible with the student. If a teacher was trained with a very different pipeline, its reasoning / output distribution may differ substantially from the student's. Because MOPD scores student-generated trajectories, those trajectories can then be out-of-distribution for the teacher—making its token-level supervision much less useful. This is primarily a problem for the consolidation setup rather than recovery, because the recovery setup by definition draws teacher models from the same training lineage as the student. One solution to this problem is MOPD warmup. Namely, before distillation, we can run a short SFT stage on a diverse set of trajectories sampled from the teacher models. This moves the student closer to teachers’ behavior, making student-generated rollouts more in-distribution. So, MOPD is not about training / distilling experts into a student. Ensuring that students / teachers are compatible is an important part of getting MOPD to work.
    @teortaxesTexDSV4.1 paper says similar things

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @cwolferesearchOne of the most interesting details I learned from the Nemotron tech reports was that certain teachers are not compatible to use in MOPD together... Multi-Teacher On-Policy Distillation (MOPD) is a distillation technique that trains a student model on its own rollouts. Given a prompt, we: - Generate a completion from the student. - Route the prompt + completion to an expert / teacher model. - Compute token-level log probabilities from the teacher. Then, we can train the student model using a reverse-KL objective (usually a sampled approximation) between student and teacher probabilities to distill the capabilities of the teacher(s) back into the student policy. Why is this useful? There are a few different setups that can be used for MOPD that differ in how they define / create the teacher models: (1) Consolidation: if we have domain-specific specialist models that have each been created using their own training pipeline, we can consolidate their advanced capabilities into the student with MOPD. This is a great way to improve the capabilities across a broad range of domains, and specialist training is independent across domains (as opposed to multi-domain RL). (2) Recovery: if we are running a sequential / multi-stage training process, we can use checkpoints from earlier training stages as teachers to restore capabilities that regress as we continue to progress through the training process. For either of these cases, we learn from Nemotron 3 Ultra that we cannot just use an arbitrary set of models as teachers for MOPD. See the following quote: “One key finding from our MOPD trials is that teacher models trained with substantially different training pipelines cannot be effectively combined through a straightforward MOPD merge, resulting in suboptimal performance.” - Nemotron 3 Ultra Tech Report MOPD teachers need to be sufficiently compatible with the student. If a teacher was trained with a very different pipeline, its reasoning / output distribution may differ substantially from the student's. Because MOPD scores student-generated trajectories, those trajectories can then be out-of-distribution for the teacher—making its token-level supervision much less useful. This is primarily a problem for the consolidation setup rather than recovery, because the recovery setup by definition draws teacher models from the same training lineage as the student. One solution to this problem is MOPD warmup. Namely, before distillation, we can run a short SFT stage on a diverse set of trajectories sampled from the teacher models. This moves the student closer to teachers’ behavior, making student-generated rollouts more in-distribution. So, MOPD is not about training / distilling experts into a student. Ensuring that students / teachers are compatible is an important part of getting MOPD to work.
    @teortaxesTexDSV4.1 paper says similar things