Token-efficiency paper selected for a NeurIPS 2026 oral presentation
A coauthor argues that longer reasoning can hurt performance and that models should learn to communicate concisely and use only necessary reasoning much earlier in training.
TLDR
A coauthor announced on September 24 that “Quantized Reasoning Models Think They Need to Think Longer, but They Do Not” was selected for a NeurIPS 2026 oral presentation, calling it a “top 0.3%” selection. The author says the work focuses on interventions during inference, when models generate answers. They argue that many tokens go toward unnecessary filler, verbose prose or poor planning that a length penalty alone cannot address.
Combined views
18
1 Source, first seen 15h ago
Token-efficiency paper selected for a NeurIPS 2026 oral presentation
A coauthor argues that longer reasoning can hurt performance and that models should learn to communicate concisely and use only necessary reasoning much earlier in training.