Reaction
A Colab for 'Planning to Learn' reportedly puts horizon loss ahead of exact PG and cross-entropy
The notebook's author says it runs in about three minutes on a free T4 GPU.
TLDR
The author reports that exact PG loses to cross-entropy even on accuracy, while horizon loss beats both. Planning CE also beats CE on its own NLL metric, they say. A reply says the work sounds similar to Schmidhuber's work from 1997 and 2011, citing 'Planning to Be Surprised' and 'RL with Self-Modifying Policies.'
Combined views
872
3 Sources, first seen ago
7 likes1 comments3 saves1 reposts