Report
AI training said to use 186 million query-document pairs while excluding benchmark-linked datasets
The author says training spans 594 datasets in 46 languages, mixing text-to-text, text-to-image and text-to-visual-document pairs.
TLDR
The author says their team trains on 186 million query-document pairs from 594 datasets in 46 languages, covering text-to-text, text-to-image and text-to-visual-document pairs. They say they removed every dataset associated with the benchmarks they evaluate on, giving up some good data for what they consider more trustworthy numbers.
Combined views
90
1 Source, first seen ago