Announcement
Marin’s training run is reportedly within 0.3% of its loss forecast at the halfway point
A post says custom kernels let Marin expand the model from 360 billion to 535 billion parameters while speeding up token processing.
TLDR
An October 1 post says Marin’s live-streamed run is halfway through training a 535-billion-parameter mixture-of-experts model on a planned 18 trillion tokens. Its evaluation loss is reportedly within 0.3% of a publicly pre-registered forecast, despite a 300-fold extrapolation. The post says the run was originally planned for 360 billion parameters, but custom kernels allowed a larger model and faster token processing.
Combined views
10.7K
2 Sources, first seen 5h ago
