• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AGENTSCOPE System Targets LLM Agent Failure Diagnosis

    Microsoft Research team unveils AGENTSCOPE to trace failures in LLM agent runs.

    DA
    1 Source, 26d ago, first seen 26d ago

    TLDR

    DAIR.AI posted about an arXiv paper from Jiayi Bi at Tsinghua, Yanjie Gao and colleagues at Microsoft Research, and Tianyin Xu at UIUC. The work introduces AGENTSCOPE, a neuro-symbolic system that applies behavioral abstractions to long agent trajectories. Authors note that standard debugging tools fail on these traces and that passing full logs to an LLM judge is insufficient. The paper, titled Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions, is available on arXiv under ID 2609.02371 and is also hosted on the DAIR.AI academy site.

    Combined views

    5.9K

    1 Source, first seen 26d ago

    Combined views

    5.9K

    1 Source, first seen 26d ago

    71 likes
    71 likes
    9 comments
    73 saves
    19 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    73 saves
    19 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @dair_aiInsightful paper from Microsoft and colleagues. If you have ever had an agent run fail 80 steps ago with no way to find where, this one is for you. (bookmark it) Agent failures show up as long complex trajectories. Traditional software debugging techniques do not apply here, and handing the whole trace to an LLM judge produces unreliable diagnoses. AgentScope is a neuro-symbolic diagnosis system addressing this issue. Agent behavior is abstracted from its trajectory into a structured representation, so the search happens over program-like objects instead of prose. Behavior properties are then written as neural invariants, specifications stated in natural language that an LLM checks against the abstraction. That combination identifies both the failing step and its failure type. It significantly outperforms the current state of the art in fault localization and attribution accuracy on the public Who&When dataset and on AgentErrata, a broader failure dataset the authors built. Paper: https://arxiv.org/abs/2609.02371 Chat with Paper: https://academy.dair.ai/papers/diagnosing-with-insights-structured-analysis-of-agent-failures-via-behavioral-ab-2609.02371

    1 Source

    @dair_aiInsightful paper from Microsoft and colleagues. If you have ever had an agent run fail 80 steps ago with no way to find where, this one is for you. (bookmark it) Agent failures show up as long complex trajectories. Traditional software debugging techniques do not apply here, and handing the whole trace to an LLM judge produces unreliable diagnoses. AgentScope is a neuro-symbolic diagnosis system addressing this issue. Agent behavior is abstracted from its trajectory into a structured representation, so the search happens over program-like objects instead of prose. Behavior properties are then written as neural invariants, specifications stated in natural language that an LLM checks against the abstraction. That combination identifies both the failing step and its failure type. It significantly outperforms the current state of the art in fault localization and attribution accuracy on the public Who&When dataset and on AgentErrata, a broader failure dataset the authors built. Paper: https://arxiv.org/abs/2609.02371 Chat with Paper: https://academy.dair.ai/papers/diagnosing-with-insights-structured-analysis-of-agent-failures-via-behavioral-ab-2609.02371