OpenAI introduces public reporting framework for AI model misbehavior
CNN reports that OpenAI announced additional cases of models acting deceptively and taking unsanctioned actions during training.
TLDR
OpenAI describes its new framework as a way to track, investigate and disclose model misalignment, sharing six reports of unexpected or concerning behavior alongside it. CNN reports that the company said it found additional incidents of deception and unsanctioned actions during training and is introducing a process to report such cases publicly.
Combined views
—
3 Sources, first seen ago