• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

SplitJEPA proposes separating stable and changing factors in AI's learned representations

A researcher reports a 71.5% success rate in a robot-arm simulation with an unseen 10-degree camera rotation.

2 Sources, 8h ago, first seen 8h ago

TLDR

A researcher behind SplitJEPA says it builds on LeJEPA by training on observation pairs that share stable factors but differ in others. They say their proof depends on sufficient variation in the changing factors across those pairs. In the PushCube robot-arm simulation, they report behavior-cloning success under an unseen 10-degree camera rotation of 71.5% with SplitJEPA, versus 12.8% with raw pixels and 48.0% with frozen R3M.

Combined views

—

2 Sources, first seen 8h ago

— likes— comments— saves— reposts

Combined views

—

2 Sources, first seen 8h ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

2 Sources

Zhuokai Zhao@zhuokaizJEPA is one of the best ideas in how machines can learn about the world, however, it is still under-explored which parts of the state it learns stay the same and which parts change. For example, a coffee mug looks different in bright morning sunlight and in dim light at night, but my hand still reaches for the same spot, because my brain keeps where the mug is separate from the lighting, which changes but does not matter for reaching. Our new paper, SplitJEPA, gives JEPA this same separation in the state it learns. In fact, the LeJEPA line of work from @ylecun, @randall_balestr, and @klindt_david already proves something strong here. Under stationary Gaussian predictive dynamics, they show that latent prediction with Gaussian regularization recovers the complete latent state up to a global orthogonal transformation, with no reconstruction needed at all. That remaining transformation is where our paper begins. Picture a map where every distance between cities is correct but there is no compass. Technically, nothing is missing, yet you still cannot tell which way is north. For LeJEPA, any orthogonal transformation of the learned state is an equally optimal solution, so each dimension of the learned representation can be a linear combination of invariant factors, which are shared across related observations, and variant factors, which change between them. SplitJEPA adds that compass. It divides the representation into an invariant block and a variant block, and alongside the predictive pairs it also trains on invariant pairs, where the two observations in a pair share the invariant factors and differ in the variant factors, such as the same physical state rendered from two cameras. An agreement loss requires the invariant block to match across each pair, which removes its linear dependence on the variant factors. Because the full transformation is orthogonal, once the invariant block holds exactly the invariant factors, the variant block is left with exactly the variant factors. The key condition is sufficient variation. SplitJEPA requires the second-moment matrix of the variant changes across pairs to be full rank, meaning no variant factor, and no combination of them, remains unchanged across all pairs. A single pair does not need to change every factor though, since the condition is on the aggregate coverage. We prove that under these conditions the invariant and variant subspaces are identified up to independent block-wise isometries, without introducing an observation decoder. The result is also stable. When the agreement holds only approximately, the leakage of variant factors into the invariant block is bounded in terms of the consistency error and the smallest eigenvalue of the variation matrix. We evaluate SplitJEPA on PushCube, a ManiSkill simulation task where a robot arm must push a cube into a goal region on a table. When the test camera is rotated 10 degrees around the scene (a camera pose not seen in training), behavior-cloning policies succeed only 12.8% of the time on raw pixels, and 48.0% on frozen R3M (a pre-trained visual representation for robot manipulation), while SplitJEPA achieves a 71.5% success rate. LeJEPA showed that the latent state can be recovered without reconstruction, and SplitJEPA extends this to show that its invariant-variant organization can, and should, be recovered the same way.8h
murat 🍥@mayferRT @zhuokaiz: JEPA is one of the best ideas in how machines can learn about the world, however, it is still under-explored which parts of t…48m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    2 Sources

    Zhuokai Zhao@zhuokaizJEPA is one of the best ideas in how machines can learn about the world, however, it is still under-explored which parts of the state it learns stay the same and which parts change. For example, a coffee mug looks different in bright morning sunlight and in dim light at night, but my hand still reaches for the same spot, because my brain keeps where the mug is separate from the lighting, which changes but does not matter for reaching. Our new paper, SplitJEPA, gives JEPA this same separation in the state it learns. In fact, the LeJEPA line of work from @ylecun, @randall_balestr, and @klindt_david already proves something strong here. Under stationary Gaussian predictive dynamics, they show that latent prediction with Gaussian regularization recovers the complete latent state up to a global orthogonal transformation, with no reconstruction needed at all. That remaining transformation is where our paper begins. Picture a map where every distance between cities is correct but there is no compass. Technically, nothing is missing, yet you still cannot tell which way is north. For LeJEPA, any orthogonal transformation of the learned state is an equally optimal solution, so each dimension of the learned representation can be a linear combination of invariant factors, which are shared across related observations, and variant factors, which change between them. SplitJEPA adds that compass. It divides the representation into an invariant block and a variant block, and alongside the predictive pairs it also trains on invariant pairs, where the two observations in a pair share the invariant factors and differ in the variant factors, such as the same physical state rendered from two cameras. An agreement loss requires the invariant block to match across each pair, which removes its linear dependence on the variant factors. Because the full transformation is orthogonal, once the invariant block holds exactly the invariant factors, the variant block is left with exactly the variant factors. The key condition is sufficient variation. SplitJEPA requires the second-moment matrix of the variant changes across pairs to be full rank, meaning no variant factor, and no combination of them, remains unchanged across all pairs. A single pair does not need to change every factor though, since the condition is on the aggregate coverage. We prove that under these conditions the invariant and variant subspaces are identified up to independent block-wise isometries, without introducing an observation decoder. The result is also stable. When the agreement holds only approximately, the leakage of variant factors into the invariant block is bounded in terms of the consistency error and the smallest eigenvalue of the variation matrix. We evaluate SplitJEPA on PushCube, a ManiSkill simulation task where a robot arm must push a cube into a goal region on a table. When the test camera is rotated 10 degrees around the scene (a camera pose not seen in training), behavior-cloning policies succeed only 12.8% of the time on raw pixels, and 48.0% on frozen R3M (a pre-trained visual representation for robot manipulation), while SplitJEPA achieves a 71.5% success rate. LeJEPA showed that the latent state can be recovered without reconstruction, and SplitJEPA extends this to show that its invariant-variant organization can, and should, be recovered the same way.8h
    murat 🍥@mayferRT @zhuokaiz: JEPA is one of the best ideas in how machines can learn about the world, however, it is still under-explored which parts of t…48m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet