A challenge to the autoregressive-error argument
A user shares a critique of treating text generated so far as both the computation and its entire state, then says parametric continual learning is unnecessary in practice.
TLDR
A user quotes a critique of the “autoregressive-error argument.” The quoted passage says that argument assumes the text generated so far is both the computation and the computation’s entire state. In a follow-up reply, the same user says Astra considers parametric continual learning—ongoing updates to a model’s learned parameters—unnecessary, and agrees that it is not needed in practice.
A challenge to the autoregressive-error argument
A user shares a critique of treating text generated so far as both the computation and its entire state, then says parametric continual learning is unnecessary in practice.
TLDR
A user quotes a critique of the “autoregressive-error argument.” The quoted passage says that argument assumes the text generated so far is both the computation and the computation’s entire state. In a follow-up reply, the same user says Astra considers parametric continual learning—ongoing updates to a model’s learned parameters—unnecessary, and agrees that it is not needed in practice.
