• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Self-generated feedback reportedly helps an AI agent solve tasks missed in 128 attempts

    The researchers say their RLTL;DR method has an agent write brief insights after failures and learn to use them.

    AG
    1 Source, 12h ago, first seen 12h ago

    TLDR

    The researchers behind RLTL;DR say their method lets an AI agent turn failed attempts into short insights and internalize what it learns. On coding and tool-calling tasks where the original policy failed all 128 attempts, they report 12–13% Pass@1 without insights provided at evaluation. Standard GRPO training stayed at 0–1% Pass@1.

    Combined views

    35

    1 Source, first seen 12h ago

    Combined views

    35

    1 Source, first seen 12h ago

    50 reposts
    50 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @alexgraveleyRT @mkirchhof_: We let an agent self-improve by exploring and internalizing ”on this sort of task, keep this sort of thing in mind“ bits of…12h

    1 Source

    @alexgraveleyRT @mkirchhof_: We let an agent self-improve by exploring and internalizing ”on this sort of task, keep this sort of thing in mind“ bits of…12h