Reaction
Exact policy gradient is claimed to trail cross-entropy on ImageNet
A user cites 4% for policy gradient versus 62% for cross-entropy on ImageNet, despite an exact gradient.
TLDR
A user argues that policy gradient can perform poorly even in image classification, where they say exploration, credit assignment and sampling noise are not factors. They cite 4% for policy gradient versus 62% for cross-entropy on ImageNet, describe a “horizon loss” tied to the remaining training budget, and claim a planning loss beats both approaches.
Combined views
10.8K
4 Sources, first seen ago
