I trained a few small run (2x GPU) MaMMUT2 weights on CC12M to validate some of the possible architecture variations, compare to similar CLIP runs
Also scaling laws paper that covers original openMaMMUT explorations, comparison to CLIP on DataComp, DFN, ReLAION