Reported model gains in 30 reinforcement-learning steps
A commenter calls the improvement dramatic and says it took far fewer training steps than they expected.
TLDR
A commenter says a model improved dramatically in 30 reinforcement-learning steps, surprising them with how few steps it needed. They wonder why training didn’t continue, saying they see “no sign of a plateau.”
Reported model gains in 30 reinforcement-learning steps
A commenter calls the improvement dramatic and says it took far fewer training steps than they expected.
TLDR
A commenter says a model improved dramatically in 30 reinforcement-learning steps, surprising them with how few steps it needed. They wonder why training didn’t continue, saying they see “no sign of a plateau.”
