• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Mamba–transformer hybrids and the role of attention

    A user argues that Mamba’s hidden states encode position implicitly, letting hybrid Mamba–transformer models dispense with explicit positional embeddings such as RoPE.

    Aleksa Gordić (水平问题)AG
    1 Source, 20d ago, first seen 20d ago

    TLDR

    A user discussing Nemotron’s architectural work says Mamba–transformer hybrids can skip explicit positional embeddings because Mamba’s hidden states already encode position. They identify two apparent weak spots in pure Mamba: copying and learning from examples in the prompt, plus long-context reasoning. On standard MMLU, they cite a 1.45-point accuracy gain when moving from no examples to five, compared with 4.38 points for the transformer. Adding a surprisingly small number of attention layers seems to address both weak spots, they argue.

    Combined views

    8.4K

    1 Source, first seen 20d ago

    Combined views

    8.4K

    1 Source, first seen 20d ago

    136 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    136 likes
    3 comments
    56 saves
    9 reposts
    3 comments
    56 saves
    9 reposts

    1 Source

    Aleksa Gordić (水平问题)@gordic_aleksathe nemotron team has been doing really cool architectural work over the past few years (scaling SSMs / mamba, hybrid models, LatentMoE, etc.) one thing that recently surprised me (but feels more obvious once you give it deeper thought) is that the hybrid SSM (mamba)-transformer architecture obviates the need for explicit positional embeddings (a la RoPE); as mamba hidden states already implicitly encode the positional information the weak parts of pure mamba seem to be: 1. copying / in-context learning abilities (e.g. it sees an accuracy improvement on standard MMLU of only 1.45 points when moving from 0 to 5 shot, compared with 4.38 for the transformer) 2. long context reasoning (consequence of recurrent state excessive compression?) and adding a surprisingly small number of attention layers seems to patch both20d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    Aleksa Gordić (水平问题)@gordic_aleksathe nemotron team has been doing really cool architectural work over the past few years (scaling SSMs / mamba, hybrid models, LatentMoE, etc.) one thing that recently surprised me (but feels more obvious once you give it deeper thought) is that the hybrid SSM (mamba)-transformer architecture obviates the need for explicit positional embeddings (a la RoPE); as mamba hidden states already implicitly encode the positional information the weak parts of pure mamba seem to be: 1. copying / in-context learning abilities (e.g. it sees an accuracy improvement on standard MMLU of only 1.45 points when moving from 0 to 5 shot, compared with 4.38 for the transformer) 2. long context reasoning (consequence of recurrent state excessive compression?) and adding a surprisingly small number of attention layers seems to patch both20d