Podcast Explores Olmo 3 Post-Training and DPO Details
Discussion covers research ideas transitioning into frontier models along with data and sequencing issues in DPO.
Nathan Lambert of AI2 announced a podcast and lecture with Scott Geng examining the practical realities of post-training the Olmo 3 model via direct preference optimization. The conversation addresses what is required for a research idea to reach a near-frontier model, the frequently messy data aspects of DPO, and organizational sequencing challenges that arise when multiple teams collaborate on open models. The session is presented as a rare public case study of these post-training workflows.
New podcast/lecture combo -- a case study in the messy details of Olmo 3 post training & DPO with @scottgeng00. It's rare to make time for these discussions, but we cover: What it takes for a research idea to make it into a (near) frontier model. The messy side of DPO (usually…

