• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    NetHack AI user asks how a reported 19.5k score maps to a progress metric

    A reply says the same method produced both high-depth and high-score AI policies, or gameplay strategies. It asks to test the depth-focused ones with the proposed metric, while questioning whether that metric might favor human-like strategies.

    JS
    1 Source, 16d ago, first seen 16d ago

    TLDR

    A user praises a reported 19.5k NetHack score using random classes and regularly reaching The Castle, then asks how that performance translates to balrogai.com’s progress metric. In reply, another user says their group wanted to demonstrate a large score improvement first because score was used in the original competition. They report having high-depth and high-score policies from different sweeps using the same method, and ask for the progress metric to be run on the high-depth policies. Their concern: whether the metric might be biased toward human-like strategies.

    Combined views

    793

    1 Source, first seen 16d ago

    Combined views

    793

    1 Source, first seen 16d ago

    10 likes
    10 likes
    1 comments
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @jsuarezOh... see, I never read the full paper because I don't follow LLM lit. We were looking for a progress metric like this. We wanted to demonstrate a huge increment on score since that was used in the original competition, before claiming that the metric needed to be changes. So we have both high-depth and high-score policies from different sweeps, same method. @finlay_sanders can you run their metric on the depth pols? The only thing I am not sure about is whether there will be a bias towards human-like strategies

    1 Source

    @jsuarezOh... see, I never read the full paper because I don't follow LLM lit. We were looking for a progress metric like this. We wanted to demonstrate a huge increment on score since that was used in the original competition, before claiming that the metric needed to be changes. So we have both high-depth and high-score policies from different sweeps, same method. @finlay_sanders can you run their metric on the depth pols? The only thing I am not sure about is whether there will be a bias towards human-like strategies