Nathan Lambert Releases RLHF Book on LLM Alignment and Post-Training
The new release pairs a print and free web book with a companion course, code, and model examples.
Entities: Nathan Lambert, Reinforcement Learning from Human Feedback
Nathan Lambert says his book *Reinforcement Learning from Human Feedback* is finished, pitching it as the resource he wanted when learning to fine-tune, align and post-train language models after ChatGPT. The project now lives at rlhfbook.com, where the site describes a free online book and course on RLHF and post-training. Lambert also linked a course page, Manning print edition and Amazon listing.
Combined views
58.3K