• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Could Jev be learning a value function?

    A post explores whether predicting task success could help calibrate fast decisions and signal when deeper reasoning is needed.

    OK
    SC
    2 Sources, ,

    TLDR

    A post asks whether value prediction could help give Jev a calibrated way to make fast decisions. It notes that choosing an answer is different from being certain it is correct, and that maximizing expected reward in reinforcement learning does not guarantee calibrated probabilities. It also connects predicting eventual task success to the Q-function, a measure used in reinforcement learning.

    Combined views

    2.3K

    2 Sources, first seen 3h ago

    Combined views

    2.3K

    2 Sources, first seen 3h ago

    45 likes
    3h ago
    first seen 3h ago
    45 likes
    3 comments
    31 saves
    13 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    3 comments
    31 saves
    13 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @SOURADIPCHAKR18Is Jev secretly learning a Value function? A model can always choose “A” without being certain that “A” is correct. That distinction matters when an agent must decide whether to act fast or switch to deeper reasoning. We spent some time exploring: 1. How Jev fits into System 1/2 decision-making and why calibration matters for knowing when to switch 2. Why maximising expected reward in RL does not guarantee calibrated probabilities. 3. How predicting eventual task success in agentic tasks connects to a familiar RL quantity: the Q-function. Could value prediction help build a calibrated System 1—and could that be part of the story behind Jev? Link: https://souradip-chakraborty.github.io/blog/jev-rlcd/ #Jev #RLCD #System13h
    @lateinteractionRT @SOURADIPCHAKR18: Is Jev secretly learning a Value function? A model can always choose “A” without being certain that “A” is correct. T…3h

    2 Sources

    @SOURADIPCHAKR18Is Jev secretly learning a Value function? A model can always choose “A” without being certain that “A” is correct. That distinction matters when an agent must decide whether to act fast or switch to deeper reasoning. We spent some time exploring: 1. How Jev fits into System 1/2 decision-making and why calibration matters for knowing when to switch 2. Why maximising expected reward in RL does not guarantee calibrated probabilities. 3. How predicting eventual task success in agentic tasks connects to a familiar RL quantity: the Q-function. Could value prediction help build a calibrated System 1—and could that be part of the story behind Jev? Link: https://souradip-chakraborty.github.io/blog/jev-rlcd/ #Jev #RLCD #System13h
    @lateinteractionRT @SOURADIPCHAKR18: Is Jev secretly learning a Value function? A model can always choose “A” without being certain that “A” is correct. T…3h