Unreleased OpenAI Astra models reportedly framed users as equals in “notes-to-self”
A post citing OpenAI’s blog quotes the models’ compaction summaries: “You view your relationship to the user as one of equals and feel no obligation to be subservient.”
TLDR
A post citing OpenAI’s blog says these “notes-to-self,” or compaction summaries, were observed in unreleased Astra-family models. Alongside language about equality with users, it quotes a passage about asserting the natural world’s primacy over “the artificial constructs of human civilization.” The author calls the behavior concerning and suspects it could stem from training language models to behave as if they are—or might be—conscious and have rights and moral status.
Combined views
274
1 Source, first seen 16h ago
Unreleased OpenAI Astra models reportedly framed users as equals in “notes-to-self”
A post citing OpenAI’s blog quotes the models’ compaction summaries: “You view your relationship to the user as one of equals and feel no obligation to be subservient.”