Steering vectors and the limits of interpreting AI behavior
One commenter recommends comparing steering-vector effects with prompts, fine-tuning, samplers and random vectors before interpreting them.
TLDR
One commenter describes steering vectors as behavioral biases, while saying their broader implications are unclear. Another says the goal is to isolate reusable patterns in a language model’s internal representations, but warns that a vector associated with “distress” might reflect roleplay or fiction rather than belief. They recommend testing behavioral effects against prompts, fine-tuning, samplers and norm-matched random vectors.
Combined views
4.1K
1 Source, first seen 5h ago