• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI

Grouping equivalent stop tokens reportedly prevents runaway output in on-policy distillation

A user sharing a paper says teacher and student models can favor different end-of-sequence tokens—the signals that end generation—even when their declared stopping sets match.

1 Source, 20d ago, first seen 20d ago

TLDR

The post describes a failure mode in on-policy distillation: a teacher model penalizes the student's preferred stopping token without transferring its own alternative, encouraging overly long responses that exhaust the generation budget. Matching decoding stop sets alone does not fix it, according to the account. The user says treating functionally equivalent end-of-sequence tokens as one shared stopping action eliminated length explosions across Qwen3, Llama and Gemma rollouts without hurting reasoning accuracy.

Combined views

—

1 Source, first seen 20d ago

— likes— comments— saves— reposts

Combined views

—

1 Source, first seen 20d ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

Taylor W. Killian@tw_killianRT @gurtej__gill_: This paper solves the problem of the Failure Mode in On-Policy Distillation. We all know how the student rollouts becom…20d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    Taylor W. Killian@tw_killianRT @gurtej__gill_: This paper solves the problem of the Failure Mode in On-Policy Distillation. We all know how the student rollouts becom…20d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet