people talk about "world modeling" in the video diffusion sense as this kind of inherent crapshoot that inherently costs 1000x more FLOPS and such in reality, almost all diffusion baselines ~everywhere are being judged on metrics like "classifier features from the Obama admin"
Users in the replies criticize video diffusion model evaluations for depending on outdated 2015 metrics such as ImageNet FID, which they see as meme-like rather than sound probabilistic references.
No Digg Deeper questions have been answered for this story yet.
Most Activity
another lesson i think that the ~entire subfield has not learned yet is that badly conditioned reparameterizations that don't buy you the simplest possible empirical wins aren't necessarily good or useful just because they have cute theory attached to them see: flow matching
~>99% of public diffusion baselines forcefully couple the idea of "the network that learns to denoise" and "the network that produces deep/nontrivial understanding". you can condition a cheap denoiser on deterministic vectors from a big model. aggressive headroom even before MoE
people only get away with this kind of thing in the diffusion literature bc the predominant reference class is a meme (imagenet FID) instead of problem constructions that you can't cheese to death via infinite proxy metric reward hacking (the kind done by academics, not agents)
people talk about "world modeling" in the video diffusion sense as this kind of inherent crapshoot that inherently costs 1000x more FLOPS and such in reality, almost all diffusion baselines ~everywhere are being judged on metrics like "classifier features from the Obama admin"
~>99% of public diffusion baselines forcefully couple the idea of "the network that learns to denoise" and "the network that produces deep/nontrivial understanding". you can condition a cheap denoiser on deterministic vectors from a big model. aggressive headroom even before MoE
people only get away with this kind of thing in the diffusion literature bc the predominant reference class is a meme (imagenet FID) instead of problem constructions that you can't cheese to death via infinite proxy metric reward hacking (the kind done by academics, not agents)
you may think: surely even if you don't do the sane thing (EDM preconditioning), the AE trained for the purposes of the diffusion model to diffuse over has normalized values instead of ones that can be arbitrarily big and unstable... surely it's finite! the answer is: not usually
for example, SD3 has bullshit heuristic weightage that covers up for the fact that vanilla flow matching has bad conditioning, and when i say bad i mean "imagine if you could get away with publishing a nanoGPT submission with no layernorm at all bc proxy evals looked good" bad
for example, SD3 has bullshit heuristic weightage that covers up for the fact that vanilla flow matching has bad conditioning, and when i say bad i mean "imagine if you could get away with publishing a nanoGPT submission with no layernorm at all bc proxy evals looked good" bad
another lesson i think that the ~entire subfield has not learned yet is that badly conditioned reparameterizations that don't buy you the simplest possible empirical wins aren't necessarily good or useful just because they have cute theory attached to them see: flow matching
unfortunately, DDPM is also not something that actually does the thing that it is supposed to do correctly when dealing with a problem that requires stronger correctness guarantees than "produces a pretty image" and needs manifold-correct samples, as shown in DIAMOND
you may think: surely even if you don't do the sane thing (EDM preconditioning), the AE trained for the purposes of the diffusion model to diffuse over has normalized values instead of ones that can be arbitrarily big and unstable... surely it's finite! the answer is: not usually
@kalomaze Inception-v3 from 2015 judging 2024 video diffusion. Everyone knows it's broken but nobody goes first.