Anthropic resumes external AI evaluations and reinforcement learning, Inside AI Policy reports
Inside AI Policy says Anthropic acknowledged an “operational security” failure. The outlet also reports criticism of its reinforcement learning techniques from advocates of pausing the technology.