Marin 535B Model Starts Open Pretraining on 18.75T Tokens
Stanford researchers launch the most transparent 535B model pretraining run to date.
TLDR
Posts from researchers including Elie Bakouch and Percy Liang announce that pretraining and midtraining have started for a 535B-A23B mixture-of-experts model named Marin. The run uses 11 GB200 NVL72 systems and will process 18.75 trillion tokens over roughly three months. Public dashboards show the exact pretraining data mixture by domain, sampled documents, live training loss, configs, and scaling laws. Participants have also pre-registered loss forecasts. The project continues Stanford CRFM's earlier open-model work that began with the 2021 Mistral effort.
Marin 535B Model Starts Open Pretraining on 18.75T Tokens
Stanford researchers launch the most transparent 535B model pretraining run to date.
