Report
Hybrid language models reportedly rely heavily on attention, even after supervised fine-tuning
The author reports that encouraging recurrent memory use improved performance, especially on long-context tasks.
TLDR
The author says recurrent–attention hybrid language models, such as Qwen3.5, rely heavily on attention even after supervised fine-tuning. They report that encouraging recurrent memory use improved performance, particularly on long-context tasks, with average gains of 4.6% on question answering and 12.1% on agentic tasks. They also report a 28.6% gain for attention-only models with multiple memory types.
Combined views
1.3K
6 Sources, first seen ago
