A proposed way to predict AI behavior: comparing agents’ hidden states
A user suggests that comparing AI agents in similar situations might also help predict when they will align, ally, or cooperate.
TLDR
A user proposes comparing an AI agent’s hidden states with those of many other agents in comparable situations to estimate the probability of repeating their outcomes. They suggest this could be a goal for interpretability research without necessarily requiring a mechanistic explanation of each agent’s behavior in isolation. The approach might also help predict alignment, alliances, and cooperation between agents. The user asks whether papers already explore this kind of work.
Combined views
2.7K
1 Source, first seen 4h ago