LLM roleplay personas keep an assistant-associated core, a post reports
The post’s author says roleplay personas in Gemma and Llama progressively differentiate from an assistant-associated core across model layers. Generated story characters, they report, lack that core.
TLDR
A post describes using sparse autoencoder (SAE) features to examine internal behavior in Gemma and Llama. The author reports that roleplay personas retain an assistant-associated core while progressively differentiating from it across model layers; generated story characters lack that core. They also report features associated with an “Immersive Simulation Mode” that distinguish immersive roleplay and story generation from the default assistant. In certain cases, the author says, those features activate even in the default assistant context, making its behavior “bizarre.”
Combined views
6.2K
1 Source, first seen 25d ago