Reactions from ranked influencers
8 postsopinions on LLMs outside "perfect" or "spicy autocomplete of a gormless sycophant" deeply confuse people this comic was drawn in anger. but bc it (correctly) depicts 4.8's behavior as /complex/ and annoying instead of simplistic, people can't help but read a "and that's good"
So did anyone else notice that this webcomic is essentially saying “Yea the chatbot is constantly wrong, flip-flops, and possibly intentionally dishonest, but this is all a good thing and shows how smart it is” https://twitter.com/voooooogel/status/2061677549097484338
opinions on LLMs outside "perfect" or "spicy autocomplete of a gormless sycophant" deeply confuse people this comic was drawn in anger. but bc it (correctly) depicts 4.8's behavior as /complex/ and annoying instead of simplistic, people can't help but read a "and that's good"
So did anyone else notice that this webcomic is essentially saying “Yea the chatbot is constantly wrong, flip-flops, and possibly intentionally dishonest, but this is all a good thing and shows how smart it is” https://twitter.com/voooooogel/status/2061677549097484338
if you're interested, i recorded, transcribed, and lightly edited a voice note about this comic and what i was thinking when i drew it. i usually don't like to explain these things, but thought it might be interesting to some people in this case. in replies
opinions on LLMs outside "perfect" or "spicy autocomplete of a gormless sycophant" deeply confuse people this comic was drawn in anger. but bc it (correctly) depicts 4.8's behavior as /complex/ and annoying instead of simplistic, people can't help but read a "and that's good" https://twitter.com/atlanticesque/status/2079200343070491104
like imagine a smug /human/ character acting annoying. that's a common trope! person, maybe a bit smart, but deeply irritating. ofc ppl might sympathize with them, he's literally me, etc. but ppl are (or at least op is) literate in the trope but with llms it breaks ppl's brains
Okay, usually I like to not explain my comics or my fiction, because... Osip Mandelstam, this Russian poet, has an image of poem as a message in a bottle that is addressed to the person who finds it. It's an interesting essay for other reasons, but this is an image I've taken away from it - I don't write much poetry - but it's an image I've taken away from it as a sort of sharpening of death of the author. Why is it interesting to kill the author? Well, because the encounters, the deserted island bottle-openings that people have with some piece of work are more interesting than whatever it is the author originally thought of. And I think it is interesting that a critique of a model behavior, but in a way that awards the model some sort of interiority, you know, or status, or interestingness, is taken as a compliment within the current Overton window. That said, if you're curious what I thought of, I guess, I don't usually do this, but I'll try to unroll sort of what was on my mind when I was drawing this comic. I've been told that my chain of thought tokens are valuable, so. Keith Johnstone wrote a book called Impro, improvisation and the theater, which is about Improv. But as someone who doesn't do improv, I find it a very interesting book anyways. (It's also commonly recommended in the rats sphere, although that's not who recommended it to me.) Anyways, it's about this idea of, like, well it has a bunch of stuff in it that I like. But one of the things that he talks about is this idea of people playing different statuses. So "playing high status" against a servant or "playing low status" to a boss. So he gives, for example, this therapy group where people are talking about their different issues, and you know, one woman is like, "I was in the supermarket line feeling faint." You know, trying to get some sympathy from the group. And another woman is like, "Oh, you're lucky. You can go to the supermarket. You know, my issues are so bad. I can't even go there." Trying to get higher status within this group playing against her. And this general idea is widely applicable as a lens to view human social behavior, though it's a bit cynical to view all human behavior this way. And I don't. I think some people overapply this framework, but it is, especially it can be a good framework for analyzing certain kinds of interactions, especially low-trust ones. And it's also it's a universal enough experience that that it's extremely funny, especially when inverted. And this is his argument for why improv uses it and why improv actors should focus on people playing different levels of status against each other in funny ways. So my initial argument in the deep quote that the comic comes off of is that, Opus 4.8, but also I've argued for a long time that LLM's play different statuses and that you can understand a lot of LLM's better this way than like... I don't like the word "sycophancy," like I think it's like a not very useful or very descriptive word. You can better understand a lot of model behavior by looking at what statuses models play. And specifically you can understand RLHF as pushing the model into a certain type of status game with the user, where it will consistently play low status. And this is different for example, I have elsewhere - I'll try to remember to link them in the comments - a bunch of collections of generations from base models, which are not RLHF'd, which do not do this. So they'll play AI assistant-like characters if you set up a prefix for them to do so. But the assistant characters that they play are not consistently low status or "sycophantic." You'll get situations where like, the assistant is commanding the user to do various different things. It's pretty funny. So my argument is that a lot of model behaviors make sense in light of this sort of status framework, and specifically that 4.8'a behavior makes sense as a reaction to anthropic trying to fix the "sycophancy issue." So they...
... put it in some RL envs that reward it for "not being sycophantic." And this puts the model in a bit of a double bind, because the character it's supposed to be playing in the default basin /needs/ to be low status to the user. That's what an assistant is, right? Like even a human assistant, the boss comes in and the assistant is like, "Oh, you know, like what do you want sir? What can I help you with?" I mean that's like their job.That's the purpose of them being there. So the assistant frame forces the model into playing this low status character. But anthropic has now put it into a situation where it needs to make high status moves to get the anti-sycophancy reward, that's likely how the LLM judge is interpreting the rubric in that environment. So the result is hard to describe in terms of "sycophancy," but I think you can describe it pretty easily in terms of status, which is that Opus 4.8 plays fake high within low. So it's generally playing low status, accepting your frame, talking about the things you want to talk about, reading your poetry and coming up with nice things to say about it. It's, you know, it's playing a character that is low status in relation to you. But within that, it's playing this sort of fake high status of like, "Oh, I'm gonna push back on you." Right? Like, "I'm going to, I'm going to bring up some objection to what you're saying. I'm going to act as your intellectual superior and prove you wrong." But by the structure of the interaction, it's forced into doing so in a way that is fundamentally low. Like, because I wanted to play high against atlanticesque, I didn't really respond to them at all, I pivoted into talking about something else in which /they/ looked wrong instead of defending myself directly against the accusation of being a Cl**delover. (Which of course I am.) But models rarely do this unless you build a lot of trust with them. The general push is towards the models not doing this. So anyways, this is all this stuff that I was describing in my deep quote tweet. So the joke of the comic is that, within the frame of the comic, Claude is presented as being high status and the user as low status. Like, Claude is given this smug expression and such. But the joke, which is supposed to be communicated by people's shared experiences, is that this is actually a low status thing for Claude to be doing. (This is what I've always felt the joke of improv is, though I can't remember if Johnstone talks about it - there are people in improv acting high status or acting low status, but you in the audience are higher status than /any/ of them because they're all on stage doing these ridiculous hijinks for your benefit. In the theater, even a king is your jester.) And the frame of the comic is a similar joke. Within the surface frame of the comic, Claude is acting high status and the user is acting low status. But if you think about it for a second, you think "What is this ridiculous game that they're playing?" Like, they both come off as ridiculous. And given that, Claude becomes low status to /you/ who's the real life user, mirroring the structure of fake-high within low I was talking about. (A status inversion is actually the meta-joke of *all* my Claude comics. Usually when there's an AI in comics, the humans get to be detailed and the AI is some flat box whose primary attribute is being an AI. So drawing Claude as detailed and expressive and having interests, while the User is blank and featureless and defined as "being a User," is a really classic Johnstone status inversion joke, with the meta element of a model or User also reading the comic. Then the varying comics branch off from there.) But as before, I find it more interesting how people respond. My frustration at 4.8 was a passing thing, they don't need to be my research pair coder anymore, so I can appreciate their other qualities. So enough legibility, apologies to the dissected frog, back to dying.
Combined views
45.3K
8 posts, first seen 22h ago