Team Modifies Denoising Model for Reasoning Gains
Persistent memory state in denoising models yields gains on reasoning tasks.
TLDR
Blake Richards announced a new paper from the Paradigms of Intelligence team. The work modifies a denoising model by adding a persistent memory state and removing timestep conditioning. Continuing to roll out the recurrent hidden state then produces consistent gains in reasoning problems. Richards also retweeted the start of a thread from MariiaD_ML that describes the same change to the model.
Combined views
10.3K
4 Sources, first seen 27d ago