DCRL post asks: What if shorter-horizon values were learned first?
A post describes DCRL as building on Transitive RL’s divide-and-conquer approach, learning shorter-horizon values before the longer-horizon values that depend on them.
TLDR
A post highlights a learning-order change in DCRL: explicitly learn shorter-horizon values first, then the longer-horizon values that depend on them. The author says DCRL builds on Transitive RL’s divide-and-conquer view of value learning and that this simple change “turns out to matter a lot.”
Combined views
9.8K
2 Sources, first seen 17d ago