Reaction
Could Jev be learning a value function?
A post explores whether predicting task success could help calibrate fast decisions and signal when deeper reasoning is needed.
TLDR
A post asks whether value prediction could help give Jev a calibrated way to make fast decisions. It notes that choosing an answer is different from being certain it is correct, and that maximizing expected reward in reinforcement learning does not guarantee calibrated probabilities. It also connects predicting eventual task success to the Q-function, a measure used in reinforcement learning.
Combined views
2.3K
2 Sources, first seen ago
