LDA-1B is a unified world model - jointly trained on 30,000 hours of human and robot data, to predict actions, forward-simulate the world in a latent representation, and predict possible futures. It's a good look at what the future of robot foundation models might look like ->
One of the more interesting innovations here is the use of a meaningful, structured DINO latent space, which captures spatial structure over merely predicting pixels