• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    “Induction heads” reportedly explain a bump in Kaplan’s AI scaling plot

    The post describes a pattern-copying mechanism: spot a token seen earlier, then copy what followed it. It says these “induction heads” can emerge in transformers with two or more layers, but not one.

    AG
    1 Source, 16d ago, first seen 16d ago

    TLDR

    A post says Anthropic researchers traced a bump in Kaplan and colleagues’ scaling-law plot to the emergence of induction heads. It describes their behavior as “[A][B] … [A] → [B]”: encountering A again prompts the head to copy the token that previously followed it. The author connects increasingly abstract versions of this behavior to in-context learning—learning from the context a model is given. According to the post, one-layer transformers do not form induction heads and therefore skip this particular phase transition, leaving their loss curve less steep; induction heads can emerge with two or more layers.

    Combined views

    20.1K

    1 Source, first seen 16d ago

    Combined views

    20.1K

    1 Source, first seen 16d ago

    318 likes
    318 likes
    4 comments
    237 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 comments
    237 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @gordic_aleksaa few years after Kaplan et al. was published, while studying induction heads*, Anthropic's mechinterp team realized that the bump in Kaplan's scaling law plot was caused by the emergence of induction heads (* an induction head is an attn head that does something like "[A][B] … [A] → [B]" i.e. when it encounters token A, it finds a previous occurence of A in the context window and copies the token that followed it. the more abstract the meanings of A and B become, the further we move from simple lexical copying and the closer we get to in-context learning) it turns out that 1-layer transformers don't form induction heads, and therefore don't go through this particular phase transition, making their loss curve less steep but with 2+ layers, induction heads can emerge

    1 Source

    @gordic_aleksaa few years after Kaplan et al. was published, while studying induction heads*, Anthropic's mechinterp team realized that the bump in Kaplan's scaling law plot was caused by the emergence of induction heads (* an induction head is an attn head that does something like "[A][B] … [A] → [B]" i.e. when it encounters token A, it finds a previous occurence of A in the context window and copies the token that followed it. the more abstract the meanings of A and B become, the further we move from simple lexical copying and the closer we get to in-context learning) it turns out that 1-layer transformers don't form induction heads, and therefore don't go through this particular phase transition, making their loss curve less steep but with 2+ layers, induction heads can emerge