Announcement
Glimmer update claims 50 tokens per second on a MacBook
A post outlining Glimmer’s first technical presentation describes distillation from Spark during training and a 131k long-context stage.
TLDR
An October 8 Glimmer update outlines its first technical presentation, including soft distillation from Spark during training, a 131k long-context stage and synthetic data for agentic and privacy use. The post says 4-bit quantization and DFlash enable 50 tokens per second on a MacBook. It also says a technical report is planned.
Combined views
2.6K
2 Sources, first seen ago
