• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology
Report

A formula aims to predict when AI chatbots might produce undesirable outputs

SecurityWeek reports that results from seven open-weight transformer models aligned with predicted immediate and delayed tipping patterns.

SecurityWeekSE
1 Source, 33m ago, first seen 33m ago

TLDR

SecurityWeek reports that George Washington University researchers developed a formula estimating how many “good” outputs a chatbot may produce before its first undesirable one. The researchers argue that accumulated conversation context can shift a model toward that tipping point. Tests across seven open-weight transformer models yielded results consistent with the predicted immediate and delayed tipping patterns. Researcher Neil Johnson says his lab has added an early-warning indicator to open-source models.

Combined views

590

1 Source, first seen 33m ago

1 likes

Combined views

590

1 Source, first seen 33m ago

1 likes

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

SecurityWeek@SecurityWeekFormula Predicts When AI Chatbots Are at Risk of Turning Bad https://www.securityweek.com/formula-predicts-when-ai-chatbots-are-at-risk-of-turning-bad/33m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    SecurityWeek@SecurityWeekFormula Predicts When AI Chatbots Are at Risk of Turning Bad https://www.securityweek.com/formula-predicts-when-ai-chatbots-are-at-risk-of-turning-bad/33m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet