Announcement
LeAVJEPA proposes a shared encoder for sound and vision
The post says dropping a modality makes audio and video learn a common representation without class labels.
TLDR
A post introduces LeAVJEPA as a self-supervised approach to learning from sound and vision. It describes one shared encoder trained without class labels and says dropping a modality makes audio and video learn a common representation.
Combined views
7.4K
2 Sources, first seen ago
