ExploreNet introduced for adaptive exploration in diffusion RL
Its creators say ExploreNet adjusts its exploration distribution to the current state, with rewards tied to rollout diversity.
TLDR
ExploreNet's creators say its state-conditioned exploration distribution is rewarded for diverse rollouts, aiming for faster, more targeted diffusion GRPO learning. They report that its noise distribution was on average 1.4 times as large as FlowGRPO's; matching the magnitude with isotropic noise suggested targeted exploration was what helped. They also say ExploreNet works with smaller groups and fewer denoising steps, and that annotators perceived greater visual changes when channels it amplified were perturbed.
Combined views
4.1K
9 Sources, first seen ago