DiPOD is claimed to boost Sudoku accuracy from 22% to 97%
The DiPOD announcement says a one-line code change boosts accuracy across reasoning tasks.
TLDR
In June, a researcher behind DiPOD called unstable post-training a major bottleneck for diffusion language models and claimed the method boosts reasoning-task accuracy, with Sudoku jumping from 22% to 97% through a one-line code change. An October 9 post said Haozhe Jiang was scheduled to present DiPOD—short for Diffusion Policy Optimization without Drifting Apart—at a COLM workshop.
Combined views
—
1 Source, first seen ago
