• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    AI training said to use 186 million query-document pairs while excluding benchmark-linked datasets

    The author says training spans 594 datasets in 46 languages, mixing text-to-text, text-to-image and text-to-visual-document pairs.

    Omar KhattabOK
    1 Source, 1h ago, first seen 1h ago

    TLDR

    The author says their team trains on 186 million query-document pairs from 594 datasets in 46 languages, covering text-to-text, text-to-image and text-to-visual-document pairs. They say they removed every dataset associated with the benchmarks they evaluate on, giving up some good data for what they consider more trustworthy numbers.

    Combined views

    90

    1 Source, first seen 1h ago

    Combined views

    90

    1 Source, first seen 1h ago

    2 reposts
    2 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Omar Khattab@lateinteractionRT @antoine_chaffin: We train on 186M query-document pairs from 594 datasets in 46 languages, mixing text-to-text, text-to-image and text-t…1h

    1 Source

    Omar Khattab@lateinteractionRT @antoine_chaffin: We train on 186M query-document pairs from 594 datasets in 46 languages, mixing text-to-text, text-to-image and text-t…1h