Positive users request a Stratechery interview with Nathan Lambert on RL, while negative users dismiss Ben Thompson's take on Chinese labs skipping strong models in RL distillation as wrong or uninformed.
Based on 4 visible X reactions from 31 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@natolambert @xeophon @benthompson I need a Statechery Interview with Nathan Lambert asap @benthompson
@natolambert @benthompson the real distillation happening here is nathan distilling ben's take down to just wrong
@natolambert @benthompson Yes love Ben, but he clearly does not understand RL. He sounded retarded
@natolambert @benthompson @benthompson Would love a @stratechery interview with Nathan on RL
Yo @benthompson I'm sorry but the Chinese labs aren't using Fable / the strongest models as teachers during RL, that's not how distillation works. It wouldn't give that big of a lift (graders during RL are messy) and you cant afford to use Fable like that.
@teortaxesTex @benthompson yes, but I think it's not "making distillation more impactful" as Ben said (the crucial point that people will latch to)
Co-sign!
Positive users request a Stratechery interview with Nathan Lambert on RL, while negative users dismiss Ben Thompson's take on Chinese labs skipping strong models in RL distillation as wrong or uninformed.
Based on 4 visible X reactions from 31 accounts; directional sample.
Ask a question below.
Published answers will appear here.