• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI

Testing a fix for inflated value estimates in reinforcement learning

A user describes a ChatGPT-assisted experiment with relative value learning, saying that centering temporal-difference errors did not improve performance on their tasks.

1 Source, 20d ago, first seen 20d ago

TLDR

One user says their Q-values—estimates of future returns—often exceed observed returns, sometimes substantially. They describe using ChatGPT to simplify relative value learning into subtracting the mean temporal-difference error from each sample’s error, but report no performance improvement on their tasks. A reply proposes a diagnostic test using cloned states: compare taking an untaken action with the highest predicted value against taking the usual action, then follow the same fixed policy and compare predicted and realized return gaps. If the predicted advantage systematically disappears, the reply argues, max-Q is feeding generalization error into its training targets.

Combined views

—

1 Source, first seen 20d ago

— likes— comments— saves— reposts

Combined views

—

1 Source, first seen 20d ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

Lucas Nestler@ClashlukeA clean test would be to clone states where an untaken action wins argmax, take that action once vs. the behavior action, follow the same frozen policy, and compare predicted Q-gap to realized return-gap. If the advantage systematically disappears, max-Q is promoting generalization error into Bellman targets.20d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    Lucas Nestler@ClashlukeA clean test would be to clone states where an untaken action wins argmax, take that action once vs. the behavior action, follow the same frozen policy, and compare predicted Q-gap to realized return-gap. If the advantage systematically disappears, max-Q is promoting generalization error into Bellman targets.20d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet