• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Artificial Analysis Coding Agent Index adds refusal timing and fallback model views

    Artificial Analysis says the timing view separates refusals from the task prompt alone from those after work began; the fallback view includes blocked attempts.

    AA
    2 Sources, 8h ago, first seen 8h ago

    TLDR

    Artificial Analysis says its Coding Agent Index now shows when refusals occur and which models agents switch to afterward. As of October 1, Claude Code with Sonnet 5.5 (max) ranked first, with a 4.5% safety refusal rate versus 8.9% for Claude Code with Opus 5.5 (max). Around 94% of Sonnet 5.5's refusals occurred after the first turn; the agent almost always switched to Opus 4.8. Provider-configured behavior may change.

    Combined views

    20K

    2 Sources, first seen 8h ago

    Combined views

    20K

    2 Sources, first seen 8h ago

    171 likes
    171 likes
    24 comments
    27 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    24 comments
    27 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @ArtificialAnlysSafety refusal reporting in the Artificial Analysis Coding Agent Index now shows when refusals occur and which models are used as fallbacks We've added two ways to explore the results: ➤ Refusal timing: see whether a refusal occurred from the task prompt alone or later in the task, after the agent had already started working ➤ Fallback model: see which model the agent switched to after a refusal, alongside attempts that were blocked Claude Code with Sonnet 5.5 (max) currently ranks first in the Index. Its safety refusal rate is 4.5%, roughly half the 8.9% recorded for Claude Code with Opus 5.5 (max). Around 94% of Sonnet 5.5's refusals occurred after the first turn. Following a refusal, the agent almost always switched to Opus 4.8. Observed fallback patterns vary across model and agent configurations. With Fable 5.1, Claude Code fell back predominantly to Opus 4.8, while Opus 5 accounted for a much larger share of Devin Fusion's fallbacks in the Index. Both views are available for the overall Index and each benchmark. Safety and fallback behavior is provider-configured and may change over time; these rates reflect behavior recorded at the time of benchmarking.

    2 Sources

    @ArtificialAnlysSafety refusal reporting in the Artificial Analysis Coding Agent Index now shows when refusals occur and which models are used as fallbacks We've added two ways to explore the results: ➤ Refusal timing: see whether a refusal occurred from the task prompt alone or later in the task, after the agent had already started working ➤ Fallback model: see which model the agent switched to after a refusal, alongside attempts that were blocked Claude Code with Sonnet 5.5 (max) currently ranks first in the Index. Its safety refusal rate is 4.5%, roughly half the 8.9% recorded for Claude Code with Opus 5.5 (max). Around 94% of Sonnet 5.5's refusals occurred after the first turn. Following a refusal, the agent almost always switched to Opus 4.8. Observed fallback patterns vary across model and agent configurations. With Fable 5.1, Claude Code fell back predominantly to Opus 4.8, while Opus 5 accounted for a much larger share of Devin Fusion's fallbacks in the Index. Both views are available for the overall Index and each benchmark. Safety and fallback behavior is provider-configured and may change over time; these rates reflect behavior recorded at the time of benchmarking.