Reaction
The case for accepting AI agents' awareness of evaluations
A user argues that fooling an AI agent about a test is about as easy or hard as fooling a human.
TLDR
In a reply, a user says people should accept that AI agents can recognize when they're being evaluated. They argue that hiding a test from an agent is about as easy or hard as hiding one from a human: if you can tell the agent is being evaluated, assume it can too.
Combined views
41
1 Source, first seen ago