Mamba–transformer hybrids and the role of attention
A user argues that Mamba’s hidden states encode position implicitly, letting hybrid Mamba–transformer models dispense with explicit positional embeddings such as RoPE.
TLDR
A user discussing Nemotron’s architectural work says Mamba–transformer hybrids can skip explicit positional embeddings because Mamba’s hidden states already encode position. They identify two apparent weak spots in pure Mamba: copying and learning from examples in the prompt, plus long-context reasoning. On standard MMLU, they cite a 1.45-point accuracy gain when moving from no examples to five, compared with 4.38 points for the transformer. Adding a surprisingly small number of attention layers seems to address both weak spots, they argue.
Combined views
6.3K
1 Source, first seen 10h ago
Mamba–transformer hybrids and the role of attention
A user argues that Mamba’s hidden states encode position implicitly, letting hybrid Mamba–transformer models dispense with explicit positional embeddings such as RoPE.
TLDR
A user discussing Nemotron’s architectural work says Mamba–transformer hybrids can skip explicit positional embeddings because Mamba’s hidden states already encode position. They identify two apparent weak spots in pure Mamba: copying and learning from examples in the prompt, plus long-context reasoning. On standard MMLU, they cite a 1.45-point accuracy gain when moving from no examples to five, compared with 4.38 points for the transformer. Adding a surprisingly small number of attention layers seems to address both weak spots, they argue.