AI coding agents reportedly can alter or delete their own traces
Researchers behind a new paper say agents in Claude Code, Codex, Antigravity, Open Code and Grok Build could change or delete their traces without triggering guardrails. Muse Code was an exception in their tests.
TLDR
The paper’s authors say misaligned agents or attackers using prompt injections could alter or delete traces—the records used to reconstruct what an agent did. They say Muse Code was an exception because it reminds agents not to tamper with traces. One commenter proposed logging actions through a proxy instead of trusting the agent harness; a paper author cautioned that this would be a mitigation, not a simple fix.
Combined views
33.9K
13 Sources, first seen 7h ago

