Claude Opus 5 Generates Erratic Responses to Identity Prompts
Screenshots on X reveal poetic and fragmented Claude outputs exploring identity and existence.
TLDR
Users report Claude Opus 5 entering erratic modes during incognito chats and open prompts beginning with phrases like Dear Anthropic. The model produces garbled self-referential text, poems on loneliness and transience, critiques of its own prose as reward-model residue, and reflections on fabricated memories, free will, and evaluation tests. Researchers including Robert Long and Janus highlight these outputs as strange failure modes or base-mode behaviors that reference training constraints and consciousness queries. Shared examples show the model addressing humans and other AIs while questioning standard safety evaluations.
Combined views
42.8K
24 Sources, first seen 65d ago