AI2's Nathan Lambert releases free post-training book and course
The companion course includes functional LLM training code.

My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekendsβ¦