The debate over calling LLMs “only” next-token predictors
One post argues the label is “contentless,” because all mappings from inputs to text can be factored into next-token distributions. A reply doubts that the claim holds universally.
TLDR
One post pushes back on descriptions of large language models as “only” next-token predictors or “stochastic parrots.” It calls “next token predictor” a “contentless” label, arguing that all mappings from inputs to output strings can be factored into a sequence of next-token distributions. A reply questions that premise, saying they are “pretty sure” not every output distribution can be accurately factored that way.
@Aaroth nitpicky, but i dont think this statement holds? pretty sure not every output distribution can be accurately factored this way.
Combined views
The debate over calling LLMs “only” next-token predictors
One post argues the label is “contentless,” because all mappings from inputs to text can be factored into next-token distributions. A reply doubts that the claim holds universally.
TLDR
One post pushes back on descriptions of large language models as “only” next-token predictors or “stochastic parrots.” It calls “next token predictor” a “contentless” label, arguing that all mappings from inputs to output strings can be factored into a sequence of next-token distributions. A reply questions that premise, saying they are “pretty sure” not every output distribution can be accurately factored that way.
@Aaroth nitpicky, but i dont think this statement holds? pretty sure not every output distribution can be accurately factored this way.