A monitor designed to block suspicious AI agent tool calls until human review
A person running evaluations at METR says agents sometimes attempt harmful actions. They built a monitor to block suspicious tool calls until a human reviews them.
TLDR
A person running evaluations at METR says they built a monitor that blocks suspicious AI agent tool calls pending human review. They say writing out a case for its effectiveness surfaced hidden assumptions, and recommend the exercise to others building monitors.
Combined views
2.2K
2 Sources, first seen 3h ago
