• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    LLMs Produce Narrower Outputs Than Training Data

    Tweet shares paper finding LLMs generate narrower answer sets than training data.

    RP
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    Rohan Paul posted that LLMs learn a narrower set of answers than their training data, even when sampling rather than using greedy decoding. The paper compares model continuations with training continuations for identical prefixes and measures output variety after accounting for the prompt. It examines OLMo, Pythia, and GPT-Neo. The tweet includes a screenshot of the academic paper as its attachment.

    Combined views

    5.1K

    2 Sources, first seen 27d ago

    Combined views

    5.1K

    2 Sources, first seen 27d ago

    14 likes
    14 likes
    4 comments
    9 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 comments
    9 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiLLMs learn a narrower set of answers than their training data, even when you sample instead of using greedy decoding. The paper compares model continuations with training continuations for the same prefixes, measuring output variety after accounting for the prompt. Across OLMo, Pythia, and GPT-Neo, training data had higher conditional diversity under every decoding method tested. For OLMo at 20K samples, the training score was 338.58, versus 218.88 with greedy decoding, 287.31 with nucleus sampling, and 297.31 with ancestral sampling. Larger models did not remove the gap, and the same pattern appeared in image generators. Simply raising temperature is not a clean fix: around T=1.9 matched OLMo's training entropy, but produced worse precision and external conditional NLL. Their alternative generates multiple candidates and reweights plausible but underused outputs. That increased measured diversity without retraining or generating new samples. – arxiv. org/abs/2609.02275 Title: "Do LLMs Capture the Diversity in their Training Data?"

    2 Sources

    @rohanpaul_aiLLMs learn a narrower set of answers than their training data, even when you sample instead of using greedy decoding. The paper compares model continuations with training continuations for the same prefixes, measuring output variety after accounting for the prompt. Across OLMo, Pythia, and GPT-Neo, training data had higher conditional diversity under every decoding method tested. For OLMo at 20K samples, the training score was 338.58, versus 218.88 with greedy decoding, 287.31 with nucleus sampling, and 297.31 with ancestral sampling. Larger models did not remove the gap, and the same pattern appeared in image generators. Simply raising temperature is not a clean fix: around T=1.9 matched OLMo's training entropy, but produced worse precision and external conditional NLL. Their alternative generates multiple candidates and reweights plausible but underused outputs. That increased measured diversity without retraining or generating new samples. – arxiv. org/abs/2609.02275 Title: "Do LLMs Capture the Diversity in their Training Data?"