i think if i was a neolab i'd probably be going all in on trying to make off policy RL happen for post training. it might never work but if it did it would be a huge deal
Benjamin Anderson proposes off-policy reinforcement learning for post-training
He acknowledges the approach carries a high failure risk.
849074K
Original post unavailable.
Sentiment
Sentiment building, check back later.
Cluster Engagement
Digg Deeper
No Digg Deeper questions have been answered for this story yet.
Posts from X
Most Activity
Most Activity
VIEWS3.6KBOOKMARKS9LIKES45REPLIES5
@andersonbcdefg how off policy are we talking
i think if i was a neolab i'd probably be going all in on trying to make off policy RL happen for post training. it might never work but if it did it would be a huge deal
@andersonbcdefg high risk, high upside research