@davidad @kaicathyc @EthanJPerez @johnschulman2 flagging for you:) definitely seems worth having someone experiment w the above given potential superalignment upsides if it works (davidad can fill in more)
@dhadfieldmenell yes! “judge mode” could also have an affordance to submit a comparison pair for human judgment, along with commentary about why it’s a tricky dilemma and what the strongest case for each side is. when human judgment is scarce, pairs could enter a tournament for…
@dhadfieldmenell the rubrics are indeed crucial in this setup in that they play a similar role as the model spec or constitution clauses. they should be drawn from a distribution that’s constantly being improved, much as (i’m given to understand) modern prosaic-alignment rubrics…
@yonashav not that i know of, that’s the question. i would do it myself, but if i use a large model i wouldn’t be able to afford enough compute to see any significant effect, and if i use a small model it won’t be smart enough to get the flywheel going 🤷