One plausible future is not enough.
A video model may generate a convincing die roll, yet produce the same few outcomes again and again.
Our new paper asks a simple question:
Does the model capture not only what can happen, but how often?
Introducing PAWBench 🧵