Peter Stone Notes TAMER as Original RLHF Algorithm
Announces availability of updated TAMER for public experimentation with reinforcement learning from human feedback.
TLDR
Peter Stone posted that TAMER was the first general-purpose RLHF algorithm and that it dates to 2008. He credited Brad Knox with releasing a new version open to anyone for experiments. The linked site states that the agent builds a model of what the user likes and then adjusts its behavior accordingly. Stone presented the update as a way for people to try the approach themselves. The packet contains only the post and site description, so the historical claim remains attributed to Stone rather than treated as independently verified.
Combined views
2.1K
4 Sources, first seen 25d ago