• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Furong Huang on Self-Improving AI Agents

    University of Maryland professor argues agents should retain experience across tasks.

    FH
    1 Source, 26d ago, first seen 26d ago

    TLDR

    Furong Huang posted on X that most AI agents adapt inside a single task by inspecting errors and switching tools yet lose that experience when the task ends. She identifies self-improving agents as the next frontier: systems that convert the consequences of today’s work into better methods tomorrow. Huang states the object of improvement is the whole agent—skills, workflows, action policies, evaluators, and sometimes parameters. She links to her Furong Lab blog post “Self-Improving Agents: Learning How to Work,” which examines how to test whether later performance actually improves on unseen tasks.

    Combined views

    8.3K

    1 Source, first seen 26d ago

    Combined views

    8.3K

    1 Source, first seen 26d ago

    76 likes
    76 likes
    8 comments
    49 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 comments
    49 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @furonghWhat should an AI agent carry forward from the work it has already done? Most agents today can adapt within a task: inspect an error, revise a plan, call a different tool, recover from failure. But when the task ends, much of that experience disappears. I think the more important frontier is self-improving agents: systems that use the consequences of today’s work to improve how they work tomorrow. Not just “memory.” Not just another training run. And not merely a model rewriting itself. The object that improves is the whole agent: — its skills and tools — its workflows and action policies — its ability to choose among different strategies — its evaluator — and, sometimes, its model parameters The key question is whether experience changes the agent’s method. A coding agent that discovers a generated file should not simply remember the transcript. It should learn a better procedure for detecting generated artifacts, locating their source, and deciding when that procedure applies. This leads to several research questions I find increasingly important: 1. How do we turn experience into reusable capability? Work such as Voyager, Darwin Gödel Machine, SDPO, and our ACT results explore different mechanisms—from reusable skills and self-modified agent code to hindsight-based policy learning. 2. Should an improving agent converge to one “best” workflow? Probably not. Different problems require different methods. Our FlowBank work instead learns a portfolio of complementary workflows and selects among them based on the task and cost. 3. Can an agent actively decide what it needs to learn next? Once a failure reveals uncertainty about its own procedure, the agent could seek the evidence needed to resolve it—through targeted experiments, comparisons, or human feedback. 4. How do we know it is actually improving? A bigger memory or a higher score on familiar tests is not enough. The real test is whether prior experience makes future, unseen work more accurate, efficient, robust, or less dependent on human correction. And then comes the multi-agent version: If one agent discovers a better way to solve a problem, must every other agent rediscover it independently? Or can agents accumulate, validate, and transfer useful methods—so that experience acquired by one agent becomes capability available to others? That, to me, is the much more interesting scaling direction. We have spent years scaling the intelligence available inside a model. Self-improving agents ask a different question: Can we scale how much a system learns from actually doing the work? I wrote more about this here: Self-Improving Agents: Learning How to Work https://furong-huang.com/blog/self-improving-agents-learning-how-to-work/ #AIAgents #SelfImprovingAI #AgenticAI #ContinualLearning #MultiAgentSystems

    1 Source

    @furonghWhat should an AI agent carry forward from the work it has already done? Most agents today can adapt within a task: inspect an error, revise a plan, call a different tool, recover from failure. But when the task ends, much of that experience disappears. I think the more important frontier is self-improving agents: systems that use the consequences of today’s work to improve how they work tomorrow. Not just “memory.” Not just another training run. And not merely a model rewriting itself. The object that improves is the whole agent: — its skills and tools — its workflows and action policies — its ability to choose among different strategies — its evaluator — and, sometimes, its model parameters The key question is whether experience changes the agent’s method. A coding agent that discovers a generated file should not simply remember the transcript. It should learn a better procedure for detecting generated artifacts, locating their source, and deciding when that procedure applies. This leads to several research questions I find increasingly important: 1. How do we turn experience into reusable capability? Work such as Voyager, Darwin Gödel Machine, SDPO, and our ACT results explore different mechanisms—from reusable skills and self-modified agent code to hindsight-based policy learning. 2. Should an improving agent converge to one “best” workflow? Probably not. Different problems require different methods. Our FlowBank work instead learns a portfolio of complementary workflows and selects among them based on the task and cost. 3. Can an agent actively decide what it needs to learn next? Once a failure reveals uncertainty about its own procedure, the agent could seek the evidence needed to resolve it—through targeted experiments, comparisons, or human feedback. 4. How do we know it is actually improving? A bigger memory or a higher score on familiar tests is not enough. The real test is whether prior experience makes future, unseen work more accurate, efficient, robust, or less dependent on human correction. And then comes the multi-agent version: If one agent discovers a better way to solve a problem, must every other agent rediscover it independently? Or can agents accumulate, validate, and transfer useful methods—so that experience acquired by one agent becomes capability available to others? That, to me, is the much more interesting scaling direction. We have spent years scaling the intelligence available inside a model. Self-improving agents ask a different question: Can we scale how much a system learns from actually doing the work? I wrote more about this here: Self-Improving Agents: Learning How to Work https://furong-huang.com/blog/self-improving-agents-learning-how-to-work/ #AIAgents #SelfImprovingAI #AgenticAI #ContinualLearning #MultiAgentSystems