• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Unpaired translation between images and text draws praise

One post says Dominik's result looked more convincing than expected; another praises novel ideas rather than just scaling.

Phillip IsolaPI
Julian TogeliusJT
8 Sources, 3h ago, first seen 3h ago

TLDR

One post calls unpaired translation between images and text a result its author had dreamed about for years, saying Dominik showed it was more possible than expected. Another says the result works and praises its skillful use of novel ideas rather than just scaling.

Combined views

19.7K

8 Sources, first seen 3h ago

458 likes13 comments320 saves39 reposts

Combined views

19.7K

8 Sources, first seen 3h ago

458 likes13 comments320 saves39 reposts
Video from X

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

8 Sources

Dominik Schnaus@dominik_schnausDINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair. It even works when the images and the captions come from different datasets. Project page: https://dominik-schnaus.github.io/unpaired-rosetta ⬇️3h
Phillip Isola@phillip_isolaThis is a result I’ve dreamt about for many years: *unpaired translation between images and text* I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:2h
Julian Togelius@togeliusCrazy that this works. But very cool. (Also, a bit of a palate cleanser to see people posting AI results that come from skillful application of novel ideas rather than just scaling.)2h
Shao-Hua Sun@shaohua0116RT @dominik_schnaus: DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a…1h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    8 Sources

    Dominik Schnaus@dominik_schnausDINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair. It even works when the images and the captions come from different datasets. Project page: https://dominik-schnaus.github.io/unpaired-rosetta ⬇️3h
    Phillip Isola@phillip_isolaThis is a result I’ve dreamt about for many years: *unpaired translation between images and text* I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:2h
    Julian Togelius@togeliusCrazy that this works. But very cool. (Also, a bit of a palate cleanser to see people posting AI results that come from skillful application of novel ideas rather than just scaling.)2h
    Shao-Hua Sun@shaohua0116RT @dominik_schnaus: DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a…1h
    Today's Rank

    #10

    Today's Rank

    #10