A formula aims to predict when AI chatbots might produce undesirable outputs
SecurityWeek reports that results from seven open-weight transformer models aligned with predicted immediate and delayed tipping patterns.
TLDR
SecurityWeek reports that George Washington University researchers developed a formula estimating how many “good” outputs a chatbot may produce before its first undesirable one. The researchers argue that accumulated conversation context can shift a model toward that tipping point. Tests across seven open-weight transformer models yielded results consistent with the predicted immediate and delayed tipping patterns. Researcher Neil Johnson says his lab has added an early-warning indicator to open-source models.
Combined views
590
1 Source, first seen ago