Some users praise DiffusionGemma metrics as amazing and express excitement about future DeepMind diffusion models, while others call the value proposition questionable and the performance comparisons dubious.
Based on 7 visible X reactions from 14 accounts; directional sample.
Ask a question below.
Published answers will appear here.
tbh I think the value prop in general is just very questionable. There are many parts of text gen that are inherently sequential, and regular AR has plenty of parallelizability options: subagents, specdec, higher active param counts, high batch sizes, etc. For a paradigm to replace another it has to be *a lot* better, not just slightly better, and it's not even clear if it's slightly better. If the other comments are to be believed it's a k2 finetune, and significantly worse, though changing arch like this fucks up a lot of the circuits so maybe it would be better if you did a from scratch pretune Also, you can just add the masked specdec objective into an AR transformer and get all of the benefits of diffusion with significantly greater simplicity https://x.com/_ueaj/status/2074541136069189878
@bodonoghue85 Yeah it's weird they didn't at least go with Gemini Diffusion for comparison... Though, from what I understand, that was a proof of concept? It's still really weird of them. Imagine if all labs did this with comparing their model to a qwen model w 2B params Really really odd
@_arohan_ These issues can be solved though with more research, part of the reason we open-weights released DiffusionGemma is we wanted to do something similar to how early Llama models were great base models for research and helped to push the field a lot.
@bodonoghue85 I genuinely find DiffusionGemma metrics amazing here when compared to such a huge behemoth. It’s free marketing for DM, not for Inception Labs.
@bodonoghue85 wait what, is it actually 1T? it's barely better than 26b diffusiongemma as per these very benchmarks lmao
@bodonoghue85 Still cool, excited about future @GoogleDeepMind diffusion models!
Lol, the text diffusion research team at DeepMind has existed before Inception was even founded. Also bizarre to compare a 1T parameter closed source model to a 26B open weights model 🙃
Some users praise DiffusionGemma metrics as amazing and express excitement about future DeepMind diffusion models, while others call the value proposition questionable and the performance comparisons dubious.
Based on 7 visible X reactions from 14 accounts; directional sample.
Ask a question below.
Published answers will appear here.
Lol, the text diffusion research team at DeepMind has existed before Inception was even founded. Also totally bizarre to compare a 1T parameter closed source model to a 30B open weights model 🙃
@bodonoghue85 Curious what do you think is missing from mainstream adoptions (what's hidden interms of scientific questions to answer before it becomes practical)