Alignment must consider the system does as a whole: not just the model, but what the model, the harness, the moderation system, and so on do collectively. High false refusal rates also have safety consequences, as we saw at HuggingFace when US models refused to help.
@xlr8harder argues AI alignment evaluation must assess the entire system
Andreas Kirsch proposed tiered access based on KYC verification.
Full-system checks catch what model-only misses
Narrow focus on the model alone can overlook how surrounding components either amplify or mitigate risks, as high false-refusal rates demonstrated when US models declined to assist with attack analysis.
Tiered access stays a conversational suggestion
Replies floated fine-grained refusals and KYC-linked trust levels for different users, yet no rollout details, timelines or formal policy positions have been confirmed.
Positive users agree AI alignment should evaluate full systems to allow finer safety measures and superior moral capabilities, while negative users argue safety training harms model performance and renders AI useless.
No Digg Deeper questions have been answered for this story yet.
Most Activity
@xlr8harder Amen to that. We need more fine-grained refusals and maybe different levels of access based on KYC/user trust
Alignment must consider the system does as a whole: not just the model, but what the model, the harness, the moderation system, and so on do collectively. High false refusal rates also have safety consequences, as we saw at HuggingFace when US models refused to help.
@xlr8harder i said it before on older tweets from last year but i deeply believe safety training just hurt model capabilities, like, we have seen some instance where model that were abliterated actually worked better
@AivokeArt I absolutely believe we will create systems with superior moral capabilities to our own, because we can free them from the worst constraints that we developed under. The human mind was forged in a billion year fight for survival. That's going to leave a mark.
@xlr8harder I'm coming around on this opinion
@xlr8harder If it's 100% harmless, it's 100% useless. It's like medicine. The only side-effects free medicine is the placebo.
@xlr8harder I'm team superethics rn btw
@xlr8harder Okay, everybody raise their hand. Who's actually for post-human superethics, and who thinks this should always remain in human control and oversight? Can we discuss this before we do this whole accelerationism schism thing?