• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Reddit Data Training Shapes LLM Response Styles

    Users observe that pre-ChatGPT Reddit answers share the explanatory tone common in LLM outputs.

    KE
    “P
    3 Sources, 62d ago, first seen 62d ago

    TLDR

    A software engineer posted that pre-ChatGPT Reddit comments often match the balanced, explanatory style of current LLMs. An academic researcher confirmed the pattern is not coincidental and traced it to three successive eras of Reddit usage in language model training. The first era began in 2019 when OpenAI filtered web pages linked from highly scored Reddit posts to create data for GPT-2. Later eras continued incorporating Reddit material, producing the persistent "Reddit voice" now familiar in model responses.

    Combined views

    533.6K

    3 Sources, first seen 62d ago

    Combined views

    533.6K

    3 Sources, first seen 62d ago

    26K likes
    26K likes
    124 comments
    1.6K saves
    805 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    124 comments
    1.6K saves
    805 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @paularamblesi love reading pre-chatgpt reddit answers and realizing just how much LLMs sound like redditors
    @ethayarajhIt's not a coincidence that LLMs have "Reddit voice". There are three eras of using Reddit for training language models: 1) Reddit as Filter (2019): In the days of GPT-2, Reddit was used as a data filter: OpenAI collected web pages linked from Reddit posts which had a score of at least 3. This became the (unreleased) WebText corpus that was used to pre-train GPT-2. 2) "Narrow" Reddit (2019-2023): Certain subreddits, such as /r/ELI5 (explain like i'm five) and /r/TL;DR (too long didn't read) were used to create narrow tasks that we wanted language models to do well on, such as question-answering and summarization. The canonical RLHF paper used data from /r/TL;DR to train reward models. 3) "Broad" Reddit (2023- onwards): The Stanford Human Preferences (SHP) dataset, created by Stanford NLP, released 385k pairwise preferences across 18 subreddits. It turns out that you can't in general assume that comment A > comment B just because A's score is higher, but you can find a subset of data for which this property is true, and this gives us a very large source of human preferences for RLHF. This is important because it's the first time that: (1) Reddit is shown to be a _broadly_ useful source of human feedback (rather than one specific subreddit for one specific task); (2) researchers start training directly on the text as opposed to using reward models, due to the rise of direct alignment methods like DPO. SHP becomes the only dataset from academia used to post-train Llama-2, the first major open-weights release for a post-trained model (the other data came from data released by OAI/Ant, collected by Meta itself, etc.). Within a year of SHP's release, Reddit started licensing its data in deals worth 200M+ dollars, and also significantly curtailed access to its data, which had historically been fairly easy to access for anyone who wanted it. At this point, it was pretty much a consensus that Reddit data was valuable, could be used broadly (for training RMs, for eval, for training directly on the text), and that frontier labs would pay a lot of money for the privilege to do so.

    3 Sources

    @paularamblesi love reading pre-chatgpt reddit answers and realizing just how much LLMs sound like redditors
    @ethayarajhIt's not a coincidence that LLMs have "Reddit voice". There are three eras of using Reddit for training language models: 1) Reddit as Filter (2019): In the days of GPT-2, Reddit was used as a data filter: OpenAI collected web pages linked from Reddit posts which had a score of at least 3. This became the (unreleased) WebText corpus that was used to pre-train GPT-2. 2) "Narrow" Reddit (2019-2023): Certain subreddits, such as /r/ELI5 (explain like i'm five) and /r/TL;DR (too long didn't read) were used to create narrow tasks that we wanted language models to do well on, such as question-answering and summarization. The canonical RLHF paper used data from /r/TL;DR to train reward models. 3) "Broad" Reddit (2023- onwards): The Stanford Human Preferences (SHP) dataset, created by Stanford NLP, released 385k pairwise preferences across 18 subreddits. It turns out that you can't in general assume that comment A > comment B just because A's score is higher, but you can find a subset of data for which this property is true, and this gives us a very large source of human preferences for RLHF. This is important because it's the first time that: (1) Reddit is shown to be a _broadly_ useful source of human feedback (rather than one specific subreddit for one specific task); (2) researchers start training directly on the text as opposed to using reward models, due to the rise of direct alignment methods like DPO. SHP becomes the only dataset from academia used to post-train Llama-2, the first major open-weights release for a post-trained model (the other data came from data released by OAI/Ant, collected by Meta itself, etc.). Within a year of SHP's release, Reddit started licensing its data in deals worth 200M+ dollars, and also significantly curtailed access to its data, which had historically been fairly easy to access for anyone who wanted it. At this point, it was pretty much a consensus that Reddit data was valuable, could be used broadly (for training RMs, for eval, for training directly on the text), and that frontier labs would pay a lot of money for the privilege to do so.