Reaction
A multimodal Jev concept for scoring how well text fits an image
A user imagines a model that scores how well each of several freeform texts fits an image. They suggest it could also give yes/no answers and produce calibrated scores.
TLDR
A user proposes a multimodal version of Jev that would take an image and a set of freeform texts, then score how well each text fits the image. They suggest it could give yes/no answers using a sigmoid and produce calibrated scores. They also speculate about releasing it with open weights and fewer than 1 billion parameters, perhaps around 400 million.
Combined views
22.7K
1 Source, first seen 12h ago
