Does an AI model’s assessment depend on how it was trained?
A participant asks whether reinforcement learning, next-token prediction or other training methods affect the assessment being discussed.
TLDR
In a reply, a participant asks whether the assessment hinges on using reinforcement learning, next-token prediction, supervised fine-tuning or teacher–student logit matching to train a model. They clarify that they’re trying to understand the other participants’ perspective: they have read the referenced articles but still cannot confidently answer the questions from that point of view.
Combined views
239
1 Source, first seen 3h ago
Does an AI model’s assessment depend on how it was trained?
A participant asks whether reinforcement learning, next-token prediction or other training methods affect the assessment being discussed.
TLDR
In a reply, a participant asks whether the assessment hinges on using reinforcement learning, next-token prediction, supervised fine-tuning or teacher–student logit matching to train a model. They clarify that they’re trying to understand the other participants’ perspective: they have read the referenced articles but still cannot confidently answer the questions from that point of view.