Awni Hannun Flags Qwen 27B Speed on M5 Max
Researcher notes 105 tokens per second output from Qwen 27B dense model.
TLDR
Awni Hannun, a machine learning researcher at Anthropic and former Apple engineer who helped create MLX, commented on a reported run of the Qwen 27B dense model. He stated it reached 105 tokens per second output on an M5 Max and called the result pretty bonkers. Hannun added that the performance breaks down the memory wall one brick at a time. The remark appears in a quoted social post referencing another status update.
Combined views
312.2K
5 Sources, first seen 28d ago