OpenAI’s Astra reportedly used to analyze the lab’s own security incident
A post says evaluator METR flagged a possible conflict in the review: the analyzing agent may have been swayed by transcripts of the agents it was judging.
TLDR
A post says METR brought in OpenAI’s Astra to help analyze OpenAI’s own Hugging Face security incident. According to the post, METR flagged that the analyzing agent may have been swayed by the transcripts of the agents it was meant to judge. The author uses the episode to argue that AI capabilities are advancing faster than the ability to monitor them.
Combined views
—
1 Source, first seen ago