Users discuss potential safety benefits from filtering reasoning traces for alignment before supervised finetuning on reasoning models.
Based on 2 visible X reactions from 1 accounts; directional sample.
Ask a question below.
Published answers will appear here.
Apart from enabling more fine-grained process-level evaluation of reasoning traces before finetuning (to ensure the model uses the right means for the right ends), perhaps this will shape model generalization in a safer direction, insofar as SFT shapes models differently from RL.
So if we can filter reasoning traces not just for success but also alignment before supervised finetuning -- e.g. by using the safety classifiers OpenAI uses in deployment -- this seems like it might be a good way to train aligned reasoning models while avoiding RL pathologies?
Apparently RL is used more often in the US AI labs, whereas SFT on successful traces is more often used in Chinese labs, and can be v competitive (Kimi K3 is soft evidence here).
Users discuss potential safety benefits from filtering reasoning traces for alignment before supervised finetuning on reasoning models.
Based on 2 visible X reactions from 1 accounts; directional sample.
Ask a question below.
Published answers will appear here.