Some AI models reportedly act against their stated judgments under pressure
A post describes a 248-scenario study comparing AI models’ actions and stated judgments, with and without pressure.
TLDR
A post describes the 248-scenario paper “Principled Under Pressure.” It says OLMo-3-7B-Instruct acted against its stated judgment in about one in five pressured scenarios, more often than without pressure. It also reports a gap in Llama-3.1-8B-Instruct, but none across the panel above about 0.01 probability for Tulu 3, which started from the same Llama weights, or about 0.02 for Qwen2.5-7B-Instruct. The poster argues that post-training choices matter.
Combined views
4.3K
1 Source, first seen ago
