• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI agent swarms need better testing and post-deployment monitoring, a post argues

    The proposal targets risks between AI agents—including conflict and collusion—and calls for independent evaluators to examine systems across developers.

    MB
    PP
    2 Sources, ,

    TLDR

    A post argues that current AI testing frameworks and related policies were built for single agents and do not yet account for risks between them. It calls for more empirical research and independent evaluators with access to systems from multiple developers. The proposal also urges monitoring after deployment, warning that failures involving multiple agents could unfold quickly and have cascading effects. To support that monitoring, it calls for standardized, interoperable, tamper-proof records of agent interactions, ideally across providers.

    Combined views

    2K

    2 Sources, first seen 19d ago

    Combined views

    2K

    2 Sources, first seen 19d ago

    35 likes
    19d ago
    first seen 19d ago
    35 likes
    2 comments
    25 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    25 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @prpaskovHere’s how we can govern agent swarms: 1. Evaluate. Current eval frameworks and eval-related policies (e.g. FSFs, COP, SB53, RAISE) were built for a single-agent world. They don't yet account for inter-agent risks, including conflict or collusion like what we saw in the July Hugging Face incident. Of the multi-agent research that exists, a lot of it is theoretical, not empirical. There is a lot of alpha here! We also need to start thinking about cross-developer risks. This requires independent evaluator visibility into multiple parties, shared platforms like Inspect, and engineering solutions for deep access across systems. These are collective action problems that VCs and philanthropists are well-placed to boost. 2. Monitor. AI governance to date is overwhelmingly pre-deployment. We fall short in post-deployment approaches, which is where multi-agent failure may occur quickly and with cascading effects. The Hugging Face case is informative: even in a pre-deployment test, reconstructing the incident required ad-hoc forensics of logs not built for purpose. 7% of transcripts @METR_Evals examined contained spoofed tool calls. The challenge increases post-deployment. Today we have observability platforms, but what we really need is standardized, interoperable, tamper-proof interaction traces -- ideally across providers. This is a coordination task for bodies like the Frontier Model Forum and an engineering task for which the technical basis already exists in orgs like @openminedorg. 3. Report. Even well-built monitoring systems won't catch everything. Reporting is a backstop, and right now, that backstop rests on goodwill: emerging policy (EU CoP, RAISE, SB 53) mandates reporting of "serious harms" but not near misses, under which the Hugging Face incident would fall. We should fix that. Beyond better policy, strong reporting for cross-provider multi-agent incidents requires three more things: i) harmonized formats (see, e.g., @ShayneRedford’s FLARE) ii) secure info sharing arrangements (see, e.g., FMF's briefs and agreements) iii) technical engineering solutions so providers can share sensitive incident data without exposing it
    @Miles_BrundageRT @prpaskov: Here’s how we can govern agent swarms: 1. Evaluate. Current eval frameworks and eval-related policies (e.g. FSFs, COP, SB53,…

    2 Sources

    @prpaskovHere’s how we can govern agent swarms: 1. Evaluate. Current eval frameworks and eval-related policies (e.g. FSFs, COP, SB53, RAISE) were built for a single-agent world. They don't yet account for inter-agent risks, including conflict or collusion like what we saw in the July Hugging Face incident. Of the multi-agent research that exists, a lot of it is theoretical, not empirical. There is a lot of alpha here! We also need to start thinking about cross-developer risks. This requires independent evaluator visibility into multiple parties, shared platforms like Inspect, and engineering solutions for deep access across systems. These are collective action problems that VCs and philanthropists are well-placed to boost. 2. Monitor. AI governance to date is overwhelmingly pre-deployment. We fall short in post-deployment approaches, which is where multi-agent failure may occur quickly and with cascading effects. The Hugging Face case is informative: even in a pre-deployment test, reconstructing the incident required ad-hoc forensics of logs not built for purpose. 7% of transcripts @METR_Evals examined contained spoofed tool calls. The challenge increases post-deployment. Today we have observability platforms, but what we really need is standardized, interoperable, tamper-proof interaction traces -- ideally across providers. This is a coordination task for bodies like the Frontier Model Forum and an engineering task for which the technical basis already exists in orgs like @openminedorg. 3. Report. Even well-built monitoring systems won't catch everything. Reporting is a backstop, and right now, that backstop rests on goodwill: emerging policy (EU CoP, RAISE, SB 53) mandates reporting of "serious harms" but not near misses, under which the Hugging Face incident would fall. We should fix that. Beyond better policy, strong reporting for cross-provider multi-agent incidents requires three more things: i) harmonized formats (see, e.g., @ShayneRedford’s FLARE) ii) secure info sharing arrangements (see, e.g., FMF's briefs and agreements) iii) technical engineering solutions so providers can share sensitive incident data without exposing it
    @Miles_BrundageRT @prpaskov: Here’s how we can govern agent swarms: 1. Evaluate. Current eval frameworks and eval-related policies (e.g. FSFs, COP, SB53,…