Nathan Lambert Shares Lecture on AI Character Training
Introduces methods for shaping model personality through constitutions and post-training.
Nathan Lambert posted that the final lecture of his course covers character training for AI models. The session examines constitutions and model specifications from labs including Anthropic and OpenAI along with practical post-training methods. He states the topic carries high real-world impact, sees extensive use at frontier labs, and has almost no empirical literature. Lambert calls the material the first long-form educational content on the subject. It is available as a video and as a chapter in his RLHF Book.
The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it: * Has potential for high real world impact * Clearly used extensively at frontier labs * Almost no empirical literature exists * More…
