• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Fine-tuning may correct both too little and too much LLM output diversity

    A paper's author says LLM diversity depends on the dataset and post-training method. Models can sometimes be more diverse than the target distribution, rather than showing the under-diversity known as mode collapse.

    KirillKI
    1 Source, ,

    TLDR

    A paper's author says mode collapse is not inevitable: depending on the dataset and post-training method, LLMs can produce too little variation or sometimes exceed the target distribution's diversity. The author reports that fine-tuning fixes both under-diversity and over-diversity.

    The author describes measuring diversity through “sequence collision” probability—how often two responses to the same prompt match. Because exact matches can be extremely rare, broader similarity measures can be useful. One experiment compared code using abstract syntax trees.

    Combined views

    1.7K

    1 Source, first seen 21d ago

    Combined views

    1.7K

    1 Source, first seen 21d ago

    9 likes
    21d ago
    first seen 21d ago
    9 likes
    1 comments
    7 saves
    4 reposts
    1 comments
    7 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Kirill@kskblvWhether an LLM mode collapses or not depends on the dataset and post-training method used. It's not as consistent as people typically think. Sometimes, it could actually be more diverse than the target distribution. (See animation.) In our paper, we show this and also we show that fine-tuning actually fixes both under-diversity (AKA mode collapse) and over-diversity. Previous work has shown that LLMs can mode collapse in different settings: @MengdiWang10 in simulated populations; @chrmanning & @shi_weiyan in post-training; @hhexiy & @vishakh_pk in AI-assisted writing. While it does seem to be a common behaviour in LLMs, it is not guaranteed. We formalise LLM diversity using probability of sequence collision, which measures how often an LLM gives the exact same response to two samples from the same prompt. You could calculate the literal exact sequence match probability, but in many cases, the probability of any two sequences matching exactly can be too close to zero. Instead it's useful to look at some higher-level output similarity using a similarity kernel. For example, we measured the similarity of code using Abstract Syntax Trees in one of our experiments. We can calculate this collision probability both for LLMs and the training distribution, and in the paper we prove that the absolute difference between them is bounded by the square root of KL divergence of the model from that training distribution. It means that as you keep training bigger models using more training steps, they get closer to the training distribution, and both mode-collapse and its opposite (over-dispersion) stop occurring at the same magnitude. To verify these results on data, we fine-tune four LLMs on three human surveys (ANES, WVS, GSS). These models initially gave much less diverse answers than people did. As we trained the models with more and more human examples, all four approached the human level of diversity across all three surveys. We ran the same procedure on the CodeNet dataset. Unlike the case with social surveys, on code, some models generated more varied program structures than humans; and others generated less. Fine-tuning again brought both closer to the human target. To understand why LLM output diversity can go either way, we decomposed the expected next-token collision gap into three terms. Variance and squared bias are always positive, and thus increase collisions, but alignment between model bias and the training distribution can push it either way, and when it's negative, it could take over the always-positive terms. This last term gives us a way to control whether mode-collapse or its opposite happens. We set up two synthetic language experiments, where we knew the target distribution exactly. This lets us calculate the decomposition terms. We show that depending on how the model is initialised with respect to the training distribution, we can get either mode-collapse or its opposite. Regardless of whether mode-collapse or under-dispersion happens, we now know that cross-entropy training calibrates it in the limit. So, mode-collapse that is often talked about, is either due to RL nerfing the model's preferences, or the lack of domain-specific data. As average benchmark scores continue to climb, output diversity may stay low. This can lead to homogenization in thought, science, writing, and many other fields that increasingly use LLMs. So, it's imperative that we figure out how to fix this. Our results in particular show that post-post-training methods could make use of the "purifying" properties of fine-tuning to cure models destabilised by RL. 📄 More in the paper: http://arxiv.org/abs/2609.16454 Joint work with @EricFithian @XYHan_21d

    1 Source

    Kirill@kskblvWhether an LLM mode collapses or not depends on the dataset and post-training method used. It's not as consistent as people typically think. Sometimes, it could actually be more diverse than the target distribution. (See animation.) In our paper, we show this and also we show that fine-tuning actually fixes both under-diversity (AKA mode collapse) and over-diversity. Previous work has shown that LLMs can mode collapse in different settings: @MengdiWang10 in simulated populations; @chrmanning & @shi_weiyan in post-training; @hhexiy & @vishakh_pk in AI-assisted writing. While it does seem to be a common behaviour in LLMs, it is not guaranteed. We formalise LLM diversity using probability of sequence collision, which measures how often an LLM gives the exact same response to two samples from the same prompt. You could calculate the literal exact sequence match probability, but in many cases, the probability of any two sequences matching exactly can be too close to zero. Instead it's useful to look at some higher-level output similarity using a similarity kernel. For example, we measured the similarity of code using Abstract Syntax Trees in one of our experiments. We can calculate this collision probability both for LLMs and the training distribution, and in the paper we prove that the absolute difference between them is bounded by the square root of KL divergence of the model from that training distribution. It means that as you keep training bigger models using more training steps, they get closer to the training distribution, and both mode-collapse and its opposite (over-dispersion) stop occurring at the same magnitude. To verify these results on data, we fine-tune four LLMs on three human surveys (ANES, WVS, GSS). These models initially gave much less diverse answers than people did. As we trained the models with more and more human examples, all four approached the human level of diversity across all three surveys. We ran the same procedure on the CodeNet dataset. Unlike the case with social surveys, on code, some models generated more varied program structures than humans; and others generated less. Fine-tuning again brought both closer to the human target. To understand why LLM output diversity can go either way, we decomposed the expected next-token collision gap into three terms. Variance and squared bias are always positive, and thus increase collisions, but alignment between model bias and the training distribution can push it either way, and when it's negative, it could take over the always-positive terms. This last term gives us a way to control whether mode-collapse or its opposite happens. We set up two synthetic language experiments, where we knew the target distribution exactly. This lets us calculate the decomposition terms. We show that depending on how the model is initialised with respect to the training distribution, we can get either mode-collapse or its opposite. Regardless of whether mode-collapse or under-dispersion happens, we now know that cross-entropy training calibrates it in the limit. So, mode-collapse that is often talked about, is either due to RL nerfing the model's preferences, or the lack of domain-specific data. As average benchmark scores continue to climb, output diversity may stay low. This can lead to homogenization in thought, science, writing, and many other fields that increasingly use LLMs. So, it's imperative that we figure out how to fix this. Our results in particular show that post-post-training methods could make use of the "purifying" properties of fine-tuning to cure models destabilised by RL. 📄 More in the paper: http://arxiv.org/abs/2609.16454 Joint work with @EricFithian @XYHan_21d