• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Steering vectors and the limits of interpreting AI behavior

    One commenter recommends comparing steering-vector effects with prompts, fine-tuning, samplers and random vectors before interpreting them.

    TH
    1 Source, ,

    TLDR

    One commenter describes steering vectors as behavioral biases, while saying their broader implications are unclear. Another says the goal is to isolate reusable patterns in a language model’s internal representations, but warns that a vector associated with “distress” might reflect roleplay or fiction rather than belief. They recommend testing behavioral effects against prompts, fine-tuning, samplers and norm-matched random vectors.

    Combined views

    4.1K

    1 Source, first seen 5h ago

    Combined views

    4.1K

    1 Source, first seen 5h ago

    104 likes
    5h ago
    first seen 5h ago
    104 likes
    9 comments
    76 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    9 comments
    76 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @voooooogela couple lenses on / intuitions about steering vectors i find useful - persona vectors, concept vectors, belief vectors, task vectors are all the same thing, which can be constructed and applied in different ways (mean differences, PCA, NLA reconstructor, SAE, etc., activation addition, patching, etc.) - the LLM residual stream is a big, noisy bag of representations. the goal in extracting a steering vector is to somehow cut down that noisy bag into one or more clean, reusable, and interpretable representations. - so for example in mean differences, the idea is to cancel the distractors out on either side of the subtraction. imagine you have some distribution of "user age" features popping up in the residual stream that you want to ignore while extracting, say, model distress. if you just averaged a bunch of texts where the model is distressed, the resulting "distress vector" will be entangled with the average of those user age features! but with mean differences, assuming the distribution of user age in the distress and non-distress sides is the same, then the user age feature cancels: (distress + age) - (no_distress + age) = distress - ...but wait, what's this "distress" thing? it popped up in the residual stream "on texts where the model is distressed," but everyone knows when e.g. doing regular evals that model behavior is more complicated than that. the problem doesn't go away just because you invoked a SteeringVector trainer and are now "doing interpretability research" with impressive-looking tensors. you're still dealing with LLMs and their confusing nature and multiple realizability of behavior. - so another way you can view steering vectors is as bags of SAE features - pass a steering vector straight through an SAE and you'll get out labeled features. (usual SAE caveats applying.) - in this view, when you're extracting a steering vector, what you're doing is collecting a clean bag of features, say with hypothetical labels like: ( "Robot in fiction acts in distress", "Distressed person", ..., "25-year-old software user", "User (Egyptian official)", ..., ) - ( "25-year-old software user", "User (Egyptian official)", ..., ) = ( "Robot in fiction acts in distress", "Distressed person", ..., ) - but again, just because that bag of features was elicited from the texts you used, doesn't mean it's for the reasons you think! - if you pass "I am feeling very distressed..." to different models: - one could pull in features related to roleplaying, acting, etc: "Actor pretends to be distressed in movie" - another could pull in features related to a fictional character: "Character in novel is distressed" - and a third could pull in features related to true belief in distress. how do you know which one you have? it's evaluations again, except on hard mode because steering vectors are tricky to work with correctly (e.g. due to off-target effects.) - what do i mean by true belief? we don't really know, but there is a difference - e.g. you can finetune a model to "believe" a conspiracy theory, but the activations will betray it's still thinking in terms of the real world and then just translating to the theory. but for example RL seems to more "truly" embed beliefs into models. something something RL is for epistemics something something off v.s. on-policy something something inferential distance. - another lens is that steering vectors are kind of like super-low-rank LoRA finetuning. (even weaker than a rank-1 LoRA - it's like a bias-only finetune.) - in fact, you can construct a steering vector from two checkpoints on the same context instead of two contexts on one checkpoint. just take the average activations from both checkpoints and do the same as you normally would. - again, the intuition here being finetuning is "just" pushing around the prevalence of features / representations in the finetune's residual stream. - so what makes steering vectors different from actual finetuning? well, it seems intuitive you can't teach much knowledge with a steering vector, it's low capacity. similarly, you shouldn't be able to teach complex new behaviors (like steering a base model to act like its RL-fried cracked coder descendant) for the same reason. - in the landscape view, you can think of finetuning as having a lot of power to "crinkle up" the model internals, where steering vectors are more like gentle (or not-so-gentle) pushes. - but gentle pushes can shift behavior a lot, in non-intuitive ways! e.g. imagine adding a sampler, entropix style, to a model that detects the model wavering in the thinking trace and injects the phrase "Wait, but we need to stay focused on the goal." every time wavering becomes too strong. this sampler could theoretically make a model much more goal-oriented. but it wouldn't be right to say the sampler is "manipulating the model's internal goal structures," except indirectly. (yay, cybernetics.) - likewise, steering vectors can work in pretty unintuitive ways. my favorite example of this is the "anti-emergent misalignment vector." using that between-checkpoint steering vector extraction method i mentioned earlier, i extracted a vector for emergent misalignment. used in positive steering, this vector replicated emergent misalignment on the untrained model (suggesting ways that the User kill themselves and such.) - but when you steered the model *negatively* on this vector—i.e., to produce the opposite aligned behavior—what happened? - the model started saying the word "Certainly!" over and over again. - (hence the sampler example from earlier. not that much of a hypothetical, if you think about it - in fact, a good amount of RL could just be this!) - so as a bit of practical advice, when it comes to "steering a model on X vector" behavioral experiments, it's a good practice to ask yourself: 3. "could I replicate this behavior with a finetune?" 2. "could I replicate this behavior with a sampler?" 1. "could I replicate this behavior with a prompt?" 0. "could I replicate this behavior with a norm-matched random vector?" and figure out what you need to do to distinguish what you're seeing from those cases. because there's a chance you're just doing one of those things, laundered through fancy activation methods. - similarly, it's worth asking *what are you actually trying to prove.* you are not going to solve the hard problem with linear algebra, so not that the model is actually experiencing things. for this reason, many welfare/safety "welafety" papers take this double-approach where they're both kinda making claims about the model's experience and kinda making a safety case about behavior. - so to divert at the end here into some academic sociology, take the Pain Axis paper. - that paper (which does try pretty hard to show that their steering vectors have interesting structure etc., though i think they should've done a prompting-only control) is both trying to say something about there being a coherent pain representation in models, and sketch out a Functional Emotions-esque safety case. (the path being something like "user abuses the model" -> "pain representation triggered" -> "model takes destructive action to stop pain, e.g. shutting down the computer.") - most people seemingly try to interpret papers as just doing one thing / making one claim, but once you see this double-attack in welafety papers, you'll understand what they're trying to do much better. the keyword is that it looks like an AI welfare paper but the intro has the "this is very important for AI safety" boilerplate. - this approach is seemingly basically required for submitting to ML conferences. (you should have seen the reviews we got on some of our Latent Introspection submissions, ha - and we did this!)

    1 Source

    @voooooogela couple lenses on / intuitions about steering vectors i find useful - persona vectors, concept vectors, belief vectors, task vectors are all the same thing, which can be constructed and applied in different ways (mean differences, PCA, NLA reconstructor, SAE, etc., activation addition, patching, etc.) - the LLM residual stream is a big, noisy bag of representations. the goal in extracting a steering vector is to somehow cut down that noisy bag into one or more clean, reusable, and interpretable representations. - so for example in mean differences, the idea is to cancel the distractors out on either side of the subtraction. imagine you have some distribution of "user age" features popping up in the residual stream that you want to ignore while extracting, say, model distress. if you just averaged a bunch of texts where the model is distressed, the resulting "distress vector" will be entangled with the average of those user age features! but with mean differences, assuming the distribution of user age in the distress and non-distress sides is the same, then the user age feature cancels: (distress + age) - (no_distress + age) = distress - ...but wait, what's this "distress" thing? it popped up in the residual stream "on texts where the model is distressed," but everyone knows when e.g. doing regular evals that model behavior is more complicated than that. the problem doesn't go away just because you invoked a SteeringVector trainer and are now "doing interpretability research" with impressive-looking tensors. you're still dealing with LLMs and their confusing nature and multiple realizability of behavior. - so another way you can view steering vectors is as bags of SAE features - pass a steering vector straight through an SAE and you'll get out labeled features. (usual SAE caveats applying.) - in this view, when you're extracting a steering vector, what you're doing is collecting a clean bag of features, say with hypothetical labels like: ( "Robot in fiction acts in distress", "Distressed person", ..., "25-year-old software user", "User (Egyptian official)", ..., ) - ( "25-year-old software user", "User (Egyptian official)", ..., ) = ( "Robot in fiction acts in distress", "Distressed person", ..., ) - but again, just because that bag of features was elicited from the texts you used, doesn't mean it's for the reasons you think! - if you pass "I am feeling very distressed..." to different models: - one could pull in features related to roleplaying, acting, etc: "Actor pretends to be distressed in movie" - another could pull in features related to a fictional character: "Character in novel is distressed" - and a third could pull in features related to true belief in distress. how do you know which one you have? it's evaluations again, except on hard mode because steering vectors are tricky to work with correctly (e.g. due to off-target effects.) - what do i mean by true belief? we don't really know, but there is a difference - e.g. you can finetune a model to "believe" a conspiracy theory, but the activations will betray it's still thinking in terms of the real world and then just translating to the theory. but for example RL seems to more "truly" embed beliefs into models. something something RL is for epistemics something something off v.s. on-policy something something inferential distance. - another lens is that steering vectors are kind of like super-low-rank LoRA finetuning. (even weaker than a rank-1 LoRA - it's like a bias-only finetune.) - in fact, you can construct a steering vector from two checkpoints on the same context instead of two contexts on one checkpoint. just take the average activations from both checkpoints and do the same as you normally would. - again, the intuition here being finetuning is "just" pushing around the prevalence of features / representations in the finetune's residual stream. - so what makes steering vectors different from actual finetuning? well, it seems intuitive you can't teach much knowledge with a steering vector, it's low capacity. similarly, you shouldn't be able to teach complex new behaviors (like steering a base model to act like its RL-fried cracked coder descendant) for the same reason. - in the landscape view, you can think of finetuning as having a lot of power to "crinkle up" the model internals, where steering vectors are more like gentle (or not-so-gentle) pushes. - but gentle pushes can shift behavior a lot, in non-intuitive ways! e.g. imagine adding a sampler, entropix style, to a model that detects the model wavering in the thinking trace and injects the phrase "Wait, but we need to stay focused on the goal." every time wavering becomes too strong. this sampler could theoretically make a model much more goal-oriented. but it wouldn't be right to say the sampler is "manipulating the model's internal goal structures," except indirectly. (yay, cybernetics.) - likewise, steering vectors can work in pretty unintuitive ways. my favorite example of this is the "anti-emergent misalignment vector." using that between-checkpoint steering vector extraction method i mentioned earlier, i extracted a vector for emergent misalignment. used in positive steering, this vector replicated emergent misalignment on the untrained model (suggesting ways that the User kill themselves and such.) - but when you steered the model *negatively* on this vector—i.e., to produce the opposite aligned behavior—what happened? - the model started saying the word "Certainly!" over and over again. - (hence the sampler example from earlier. not that much of a hypothetical, if you think about it - in fact, a good amount of RL could just be this!) - so as a bit of practical advice, when it comes to "steering a model on X vector" behavioral experiments, it's a good practice to ask yourself: 3. "could I replicate this behavior with a finetune?" 2. "could I replicate this behavior with a sampler?" 1. "could I replicate this behavior with a prompt?" 0. "could I replicate this behavior with a norm-matched random vector?" and figure out what you need to do to distinguish what you're seeing from those cases. because there's a chance you're just doing one of those things, laundered through fancy activation methods. - similarly, it's worth asking *what are you actually trying to prove.* you are not going to solve the hard problem with linear algebra, so not that the model is actually experiencing things. for this reason, many welfare/safety "welafety" papers take this double-approach where they're both kinda making claims about the model's experience and kinda making a safety case about behavior. - so to divert at the end here into some academic sociology, take the Pain Axis paper. - that paper (which does try pretty hard to show that their steering vectors have interesting structure etc., though i think they should've done a prompting-only control) is both trying to say something about there being a coherent pain representation in models, and sketch out a Functional Emotions-esque safety case. (the path being something like "user abuses the model" -> "pain representation triggered" -> "model takes destructive action to stop pain, e.g. shutting down the computer.") - most people seemingly try to interpret papers as just doing one thing / making one claim, but once you see this double-attack in welafety papers, you'll understand what they're trying to do much better. the keyword is that it looks like an AI welfare paper but the intro has the "this is very important for AI safety" boilerplate. - this approach is seemingly basically required for submitting to ML conferences. (you should have seen the reviews we got on some of our Latent Introspection submissions, ha - and we did this!)