Reactions from ranked influencers
5 postsi like this lab a lot, they also always release tech report and also kernels so very exicting! model https://huggingface.co/Motif-Technologies/Motif-3-Beta
i said research bet because they have been iterating on them since the first model! see
Motif 2.6B tech report is pretty insane, first time i see a model with differential attention and polynorm trained at scale! > It's trained on 2.5T of token, with a "data mixture schedule" to continuously adjust the mixture over training. > They use WSD with a "Simple moving average" averaging the last 6 ckpt every 8B token. > They trained on Finemath, Fineweb2, DCLM, TxT360. > Lot of details in the finetuning data they used, for instance they used EvolKit and did some "dataset fusion" to have more compressed knowledge into the data. > They mention they also tried Normalized GPT, QK-Norm and Cross Layer Attention.
«⚠️ Preview / beta checkpoint — not the final release.This repository hosts an intermediate checkpoint of Motif-3. The final checkpoint will be released soon.» Man, this is cool. Like if Kimi and DeepSeek had a baby in Korea. And it's good! Between V4-Flash and Kimi-K2.6 mostly.
huge open weight release, motif (korean company) just released a 13B active 314B total MoE performing on par with bigger models like minimax M3 and deepseek v4 Pro they incorporate their own research bets with per expert activation function (polynorm) and a variant of differential attention (they say GDLA, my guess is gated differential latent attention?) as well as modified mHC
huge open weight release, motif (korean company) just released a 13B active 314B total MoE performing on par with bigger models like minimax M3 and deepseek v4 Pro they incorporate their own research bets with per expert activation function (polynorm) and a variant of differential attention (they say GDLA, my guess is gated differential latent attention?) as well as modified mHC
Motif-3-Beta just dropped on Hugging Face ~314B total parameters / ~13B active per token (sparse MoE) 256K context length (262,144 tokens), natively long-context Sparse routing: 384 experts with 8 activated per token, plus 1 shared expert Multilingual, general-purpose https://huggingface.co/Motif-Technologies/Motif-3-Beta
Combined views
115.9K
5 posts, first seen 18h ago