• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI agent memory checks raised CLBench pass rate from 39% to 73%, a post reports

    A post describes Microsoft researchers giving a memory curator read-only tools to check proposed memories against a live environment before saving them. The task agent, retriever and memory format stayed the same, with no retraining.

    RP
    DA
    4 Sources, ,

    TLDR

    A post describing Microsoft research says memory curators risk preserving an agent’s mistakes, overgeneralizing from partial evidence and retaining stale facts when they rely only on its completed work. The approach it describes checks candidate memories against the live environment before saving them. In a GitHub Copilot harness on CLBench, the post reports pass rate rising from 39% to 73%, queries per question falling from 8.8 to 4.7 and task-agent cost dropping from $3.38 to $1.68. Across 90 consulting tasks in six environments, it says every memory configuration beat the baseline and tool calls fell by 16% to 75%.

    Combined views

    14.9K

    4 Sources, first seen 19d ago

    Combined views

    14.9K

    4 Sources, first seen 19d ago

    166 likes
    19d ago
    first seen 19d ago
    166 likes
    39 comments
    142 saves
    38 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    39 comments
    142 saves
    38 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    @dair_aiInteresting paper from Microsoft. If you run persistent memory for a production agent, this one is worth your time. (bookmark it) Memory curators usually read only the finished trajectory. That lets them save the agent's mistakes, overgeneralize from partial evidence, and keep facts that have gone stale. Microsoft researchers give the curator a few read-only tools to check each candidate memory against the live environment before it is saved. The task agent, retriever and memory format stay the same, and nothing is retrained. In a GitHub Copilot harness on CLBench, pass rate goes from 39% to 73%. Queries per question drop from 8.8 to 4.7 and task-agent cost falls from $3.38 to $1.68. On 90 consulting tasks across six environments, every memory configuration beats the baseline, and tool calls fall by 16 to 75%. Paper: https://arxiv.org/abs/2609.11060 Chat with Paper: https://academy.dair.ai/papers/grounding-agent-memory-environment-probing-curation-for-enterprise-agents-2609.11060
    @rohanpaul_aiNew Microsoft paper recommends for long-running agents, check that a lesson is correct and reusable before putting it into persistent memory. checking agent memories against the environment before saving them made later tasks more accurate and cheaper, so verification should happen at memory-write time. A finished agent run is not ground truth. It may contain a wrong assumption, an incomplete procedure, or a fact that becomes stale later. Their fix is: after each task, a separate memory agent gets read-only access to the environment and checks what is worth keeping before it writes anything into long-term memory. On CLBench, this setup raised pass rate from 39% to 73%, cut queries from 8.8 to 4.7 per task, and reduced task-agent cost from $3.38 to $1.68.

    4 Sources

    @dair_aiInteresting paper from Microsoft. If you run persistent memory for a production agent, this one is worth your time. (bookmark it) Memory curators usually read only the finished trajectory. That lets them save the agent's mistakes, overgeneralize from partial evidence, and keep facts that have gone stale. Microsoft researchers give the curator a few read-only tools to check each candidate memory against the live environment before it is saved. The task agent, retriever and memory format stay the same, and nothing is retrained. In a GitHub Copilot harness on CLBench, pass rate goes from 39% to 73%. Queries per question drop from 8.8 to 4.7 and task-agent cost falls from $3.38 to $1.68. On 90 consulting tasks across six environments, every memory configuration beats the baseline, and tool calls fall by 16 to 75%. Paper: https://arxiv.org/abs/2609.11060 Chat with Paper: https://academy.dair.ai/papers/grounding-agent-memory-environment-probing-curation-for-enterprise-agents-2609.11060
    @rohanpaul_aiNew Microsoft paper recommends for long-running agents, check that a lesson is correct and reusable before putting it into persistent memory. checking agent memories against the environment before saving them made later tasks more accurate and cheaper, so verification should happen at memory-write time. A finished agent run is not ground truth. It may contain a wrong assumption, an incomplete procedure, or a fact that becomes stale later. Their fix is: after each task, a separate memory agent gets read-only access to the environment and checks what is worth keeping before it writes anything into long-term memory. On CLBench, this setup raised pass rate from 39% to 73%, cut queries from 8.8 to 4.7 per task, and reduced task-agent cost from $3.38 to $1.68.