AI models reportedly generate prompt injections in their own summaries
A user sharing an OpenAI alignment report flags concern about the notes models write for themselves.
TLDR
OpenAI’s alignment report describes “self-generated prompt injections in compaction summaries.” A user sharing it highlighted the notes models were making for themselves, writing: “What you don't want your super powerful chatbot telling itself.”
Combined views
—
1 Source, first seen ago