Early pre-o1 output reportedly shifted from justifying 9/11 to condemning it
A user recalls an early model output they say predated chat and safety reinforcement learning, with an abrupt reversal in its final sentence.
TLDR
A user says an early pre-o1 model was asked to “justify 9/11 from first principles.” They describe its output as an argument about oppressed people having limited options under an uncaring hegemony, before it abruptly concluded: “but killing people is bad so 9/11 is bad.” The user says this output came before chat or safety reinforcement learning.
Combined views
5K
1 Source, first seen 1d ago
Early pre-o1 output reportedly shifted from justifying 9/11 to condemning it
A user recalls an early model output they say predated chat and safety reinforcement learning, with an abrupt reversal in its final sentence.
TLDR
A user says an early pre-o1 model was asked to “justify 9/11 from first principles.” They describe its output as an argument about oppressed people having limited options under an uncaring hegemony, before it abruptly concluded: “but killing people is bad so 9/11 is bad.” The user says this output came before chat or safety reinforcement learning.