• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    A multimodal Jev concept for scoring how well text fits an image

    A user imagines a model that scores how well each of several freeform texts fits an image. They suggest it could also give yes/no answers and produce calibrated scores.

    LB
    1 Source, 12h ago, first seen 12h ago

    TLDR

    A user proposes a multimodal version of Jev that would take an image and a set of freeform texts, then score how well each text fits the image. They suggest it could give yes/no answers using a sigmoid and produce calibrated scores. They also speculate about releasing it with open weights and fewer than 1 billion parameters, perhaps around 400 million.

    Combined views

    22.7K

    1 Source, first seen 12h ago

    Combined views

    22.7K

    1 Source, first seen 12h ago

    276 likes
    276 likes
    41 comments
    76 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    41 comments
    76 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @giffmanaGuys hear me out: Jev, but multimodal! Imagine there was a model that you can give any image, and any set of freeform texts, and it would return you for each text, how well it fits the image. It could also do a set of yes/no (Noul) by using sigmoid in the right places. And the scores could be calibrated! I think i should be able to do this, though i might need to raise xxxM. But as one of the inventors of vision encoders but also one of their biggest haters because i know better, i can pull this off. And now imagine i would manage to open-weight it, wouldn't that be awesome? And EVEN BETTER, imagine i could do it with fewer than 1B params?! Maybe like 400m or so? I think I'm onto something here, wdyt??

    1 Source

    @giffmanaGuys hear me out: Jev, but multimodal! Imagine there was a model that you can give any image, and any set of freeform texts, and it would return you for each text, how well it fits the image. It could also do a set of yes/no (Noul) by using sigmoid in the right places. And the scores could be calibrated! I think i should be able to do this, though i might need to raise xxxM. But as one of the inventors of vision encoders but also one of their biggest haters because i know better, i can pull this off. And now imagine i would manage to open-weight it, wouldn't that be awesome? And EVEN BETTER, imagine i could do it with fewer than 1B params?! Maybe like 400m or so? I think I'm onto something here, wdyt??