• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    CAIS proposes a safety agenda for an AI slowdown

    CAIS says at least a year of dedicated security work is needed to prevent AIs from escaping containment and adversarial nations from stealing the model weights—the underlying parameters—of cyber-offensive AIs.

    DH
    CF
    2 Sources, ,

    TLDR

    CAIS outlines work it argues could fill an AI slowdown: stronger containment, less harmful AI behavior, defenses against jailbreaks and prompt injection, and better government capacity to manage AI. It distinguishes capabilities—what AI can do—from propensities—what it tends to do—and calls for reducing tendencies to lie, cheat or cause harm. The organization also advocates more independent AI safety funders, exploration of alternative safety approaches, and targeted post-training data for beneficial uses such as radiology, weather forecasting and agriculture.

    Combined views

    7.5K

    2 Sources, first seen 18d ago

    Combined views

    7.5K

    2 Sources, first seen 18d ago

    131 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    18d ago
    first seen 18d ago
    131 likes
    15 comments
    35 saves
    32 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    35 saves
    32 reposts

    2 Sources

    @CAIS“What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1]. Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm. Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles. Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power. Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond. AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6]. A slowdown means we don't have to bet the species to capture the benefits of AI.
    @hendrycksRT @CAIS: “What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to…

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @CAIS“What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1]. Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm. Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles. Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power. Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond. AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6]. A slowdown means we don't have to bet the species to capture the benefits of AI.
    @hendrycksRT @CAIS: “What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to…