“Induction heads” reportedly explain a bump in Kaplan’s AI scaling plot
The post describes a pattern-copying mechanism: spot a token seen earlier, then copy what followed it. It says these “induction heads” can emerge in transformers with two or more layers, but not one.
TLDR
A post says Anthropic researchers traced a bump in Kaplan and colleagues’ scaling-law plot to the emergence of induction heads. It describes their behavior as “[A][B] … [A] → [B]”: encountering A again prompts the head to copy the token that previously followed it. The author connects increasingly abstract versions of this behavior to in-context learning—learning from the context a model is given. According to the post, one-layer transformers do not form induction heads and therefore skip this particular phase transition, leaving their loss curve less steep; induction heads can emerge with two or more layers.
Combined views
20.1K
1 Source, first seen 16d ago