Anthropic resumes external AI evaluations and reinforcement learning, Inside AI Policy reports
Inside AI Policy says Anthropic acknowledged an “operational security” failure. The outlet also reports criticism of its reinforcement learning techniques from advocates of pausing the technology.
TLDR
Inside AI Policy reports that Anthropic is again evaluating its AI models in third-party environments and training them using reinforcement learning. The outlet says Anthropic acknowledged an “operational security” failure and describes the training techniques as drawing criticism from proponents of a pause on the technology.
Combined views
39
1 Source, first seen 29d ago