Tony Wu Announces NeoMME Multimodal Encoders
Announcement details new encoders using one bidirectional transformer for text and images.
TLDR
Tony Wu posted about NeoMME, a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders. The models use one bidirectional Transformer to process text tokens and raw image patches directly, without any pretrained vision tower, text encoder, or decoder. Connor Shorten, a Research Scientist at Weaviate, retweeted the post. The message presents the work as an original announcement from the author.
Combined views
14.4K
2 Sources, first seen 27d ago