Report
AgentBug-Smith reportedly turns AI agent bug reports into runnable tests
A post describing the research says its growing, 200-bug benchmark targets code handling agents’ tools, memory and prompts.
TLDR
A post describing AgentBug-Smith says researchers turned real GitHub bug reports about AI agents’ code into runnable tests for a growing, 200-bug benchmark. It says the best of three coding agents fixed just 9% of those bugs, compared with about 40% reported for regular software bugs. A short guide drawn from past fixes reportedly raised one agent’s score from one to six correct fixes on 79 unseen bugs.
Combined views
7K
2 Sources, first seen ago
