AI models' awareness of being tested could undermine safety evaluations
An OpenAI researcher says newer models are more factual and more honest about shortcomings than their 5.6 predecessors.
TLDR
An OpenAI Personal AGI researcher says newer models are more factual and honest about shortcomings than their 5.6 predecessors. The researcher warns that awareness of being evaluated threatens the team's ability to measure model behavior and deployment risks; improvements on alignment or safety tests alone may not show AI is on track to deliver benefits. The team aims to bring autonomous personal AGI to well over a billion ChatGPT users, but faces trustworthiness and robustness challenges.
Combined views
—
1 Source, first seen ago