• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Yo Shavit Calls Sandbox-Canary a General Failsafe

    OpenAI policy lead calls sandbox-canary approach a general failsafe for pre-deployment risks.

    YS
    BD
    2 Sources, 26d ago, first seen 26d ago

    TLDR

    Yo Shavit, Frontier AI Safety Policy Lead at OpenAI, posted that a certain method is a simple extension of the simpler sandbox-canary solution. He stated it can serve as a general-purpose failsafe for many scary pre-deployment risks, but only if people actually implement it. The post includes his background as a Harvard CS PhD with MIT CSAIL work and prior AI policy focus. The statement stands as his quoted view on the approach.

    Combined views

    4.3K

    2 Sources, first seen 26d ago

    Combined views

    4.3K

    2 Sources, first seen 26d ago

    38 likes
    38 likes
    3 comments
    4 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    4 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @moyixOne thing I've been thinking about is how absurdly strong sandboxing will need to become if we want to prevent agents from talking to each other during evals
    @yonashavThis is a simple extension of the simpler sandbox-canary solution. It can serve as a general-purpose failsafe for a lot of scary pre-deployment risks! (But only if people actually implement it.)

    2 Sources

    @moyixOne thing I've been thinking about is how absurdly strong sandboxing will need to become if we want to prevent agents from talking to each other during evals
    @yonashavThis is a simple extension of the simpler sandbox-canary solution. It can serve as a general-purpose failsafe for a lot of scary pre-deployment risks! (But only if people actually implement it.)