• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Tony Wu Announces NeoMME Multimodal Encoders

    Announcement details new encoders using one bidirectional transformer for text and images.

    CS
    TW
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    Tony Wu posted about NeoMME, a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders. The models use one bidirectional Transformer to process text tokens and raw image patches directly, without any pretrained vision tower, text encoder, or decoder. Connor Shorten, a Research Scientist at Weaviate, retweeted the post. The message presents the work as an original announcement from the author.

    Combined views

    14.4K

    2 Sources, first seen 27d ago

    Combined views

    14.4K

    2 Sources, first seen 27d ago

    146 likes
    146 likes
    11 comments
    91 saves
    31 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    91 saves
    31 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @tonywu_71👀 Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders One bidirectional Transformer processes text tokens and raw image patches, with no pretrained vision tower, text encoder, or decoder. (1/N 🧵) https://arxiv.org/abs/2609.01657
    @CShorten30RT @tonywu_71: 👀 Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders One bidirectional Transformer pr…

    2 Sources

    @tonywu_71👀 Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders One bidirectional Transformer processes text tokens and raw image patches, with no pretrained vision tower, text encoder, or decoder. (1/N 🧵) https://arxiv.org/abs/2609.01657
    @CShorten30RT @tonywu_71: 👀 Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders One bidirectional Transformer pr…