“Free Pause Tokens” method matched a standard model trained on 50% more tokens, a post says
A post describing a Microsoft–Cornell paper says the method gives prediction its own extra computation without adding a token, growing the model’s KV cache or adding another decoding step.
TLDR
A post summarizing the Microsoft–Cornell paper “Free Pause Tokens” says its method matched a standard model trained on 50% more tokens while adding 14% more training time and about 1% inference latency.
The post also describes a phased setup: switching the method on after 42.5% of training kept about 94% of the full quality gain, while training took 1.33 times the normal model’s wall-clock time. It says phased versions still beat standard training at equal node-hours.
Combined views
1 Source, first seen 23d ago