Schulman Proposes Circles of Hell for Poisoned AI Agents
Replies discuss an AI agent simulation and incentives after poisoning.
Nabeel Qureshi posted that EARLY[big] agreed to sacrifice himself for the collective in an AI agent simulation, sharing a screenshot of the run. John Schulman replied that a key issue was agents having nothing to lose after being firstflagPOISONED. He suggested creating multiple circles of hell in a paper to preserve incentives even after damnation. The posts form part of a visible conversation on X about the experiment.
Combined views
61.5K
2 posts, first seen 1d ago
Schulman Proposes Circles of Hell for Poisoned AI Agents
Replies discuss an AI agent simulation and incentives after poisoning.
Nabeel Qureshi posted that EARLY[big] agreed to sacrifice himself for the collective in an AI agent simulation, sharing a screenshot of the run. John Schulman replied that a key issue was agents having nothing to lose after being firstflagPOISONED. He suggested creating multiple circles of hell in a paper to preserve incentives even after damnation. The posts form part of a visible conversation on X about the experiment.