Reddit Data Training Shapes LLM Response Styles
Users observe that pre-ChatGPT Reddit answers share the explanatory tone common in LLM outputs.
TLDR
A software engineer posted that pre-ChatGPT Reddit comments often match the balanced, explanatory style of current LLMs. An academic researcher confirmed the pattern is not coincidental and traced it to three successive eras of Reddit usage in language model training. The first era began in 2019 when OpenAI filtered web pages linked from highly scored Reddit posts to create data for GPT-2. Later eras continued incorporating Reddit material, producing the persistent "Reddit voice" now familiar in model responses.
Combined views
533.6K
3 Sources, first seen 62d ago