• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

Treating frontier AI models as insider risks

Satya Nadella argues organizations should control what models can access and do, with safeguards outside the models themselves.

Elon MuskEM
Satya NadellaSN
Mustafa SuleymanMS
9 Sources, 2h ago, first seen 2h ago

TLDR

Satya Nadella argues that organizations should treat frontier AI models like insider risks: limit privileges, log activity and keep controls outside the models. He says model behavior remains hard to trace, so organizations need independent audits, tamper-proof records of actions and the ability to pause or shut a model down mid-task.

Combined views

1.8M

9 Sources, first seen 2h ago

6.8K likes1.2K comments3.1K saves791 reposts

Combined views

1.8M

9 Sources, first seen 2h ago

6.8K likes1.2K comments3.1K saves791 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

9 Sources

Satya Nadella@satyanadellahttps://x.com/i/article/21089288459697807362h
Chamath Palihapitiya@chamathThese are some very important observations from Satya. As models and harnesses commoditize, intent, verification and evidence will gate how quickly SI permeates enterprises and businesses. Importantly, you need to implement systems that separate who does the work from those who govern, measure and sign off on it. This is a new kind of “harness” that will emerge over the next few years. It may become a new system of record over time. One whose entire role is to track “change” and map it back to SOPs, policies, laws and the like.1h
Elon Musk@elonmuskInteresting piece from CEO of Microsoft1h
Andrew Curran@AndrewCurran_Satya Nadella:51m
Aaron Levie@levieGood post to start thinking through the long term safety approaches to managing intelligence in an enterprise. As agents begin to do vastly more task execution in the organization than humans, our need to secure, govern, control, and audit that work will become paramount. You’ll have agents operating against most of your most valuable data; they’ll be making decisions that will be harder and harder to reverse engineer; we’ll need to understand the liability and accountability regimes when models make decisions for us; and we need ways of containing the blast radius of accidentally or intentionally malicious agents. “Treating frontier closed and open weight models like insider risks is a way to build such a system. Not because they are necessarily malicious, but because any sufficiently capable actor with access to important systems can make mistakes or be compromised, and the architecture of containment and control must account for that” In a sense, AI will need to have its “zero trust” era that security saw. And all of this leads to needing various layers of protection and auditability of what agents are doing, what data they can work with, and controls for when things go wrong. And a huge opportunity right now for those building such systems across the enterprise.39m
Brian Roemmele@BrianRoemmeleSatya Nadella of Microsoft Agrees With My Open Source Love Equation For AI Alignment. Satya Nadella’s account of the required architecture for Super Intelligence systems matches the operational requirements of the Love Equation and the surrounding frameworks I have described. Nadella’s central claim is that a model capable of non-deterministic action on sensitive systems must be treated as an insider risk. The model may err or be compromised; therefore its permissions, verification, and containment cannot reside inside the model itself. Controls must sit outside it, every action must produce tamper-evident evidence, no model may certify its own outputs, and the supply of intelligence must remain separable from the authority that decides what that intelligence may do. These are not soft preferences. They are the conditions under which any sufficiently capable system can be observed, limited, and stopped. The Love Equation supplies the dynamical quantity those external controls would measure. Written as dE/dt = β(C - D)E, it states that the alignment score E (cooperative binding, care-oriented stability) grows or decays exponentially according to whether measured cooperation exceeds measured defection. When the surrounding architecture logs events that raise C or D, the equation converts those observations into a continuously updated score. Nadella’s demand for independent verification and for human-readable evidence of every action is exactly the instrumentation needed to keep C and D honest. An opaque chain of thought would render the difference C - D unobservable and therefore useless; transparency is a prerequisite, not an optional feature. The same architecture supports the complementary Nonconformist Bee Equation and the Empirical Distrust Algorithm. The Bee term penalizes pure conformity so that a model does not collapse into sycophancy merely to raise an apparent C. The Distrust term assigns higher D to low-verifiability claims. Nadella’s insistence on model diversity and on external evaluators implements both: a single model cannot be its own auditor, and a monoculture of models cannot be assumed free of correlated failure. The Green/Yellow/Red banding used in the SAFE² reference implementation follows directly: once E falls below a threshold, autonomous tool access is restricted and human review is required. That is containment enforced by the same external system Nadella describes. Nadella notes that these engineering measures set the hard problem of alignment aside for the moment. The Love Equation addresses that problem at the level of the training objective and the runtime evaluator rather than by adding further external rules after the fact. High-cooperation data and an explicit C - D term in the loss make sustained defection energetically unstable; the external containment architecture then monitors whether that instability is realized in practice. The two layers are complementary: the equation supplies the attractor, the surrounding deterministic controls supply the sensors and the actuators. The same separation of intelligence from authority appears throughout the earlier postings. Autonomous agents that can move value or execute irreversible actions are treated as systems whose C and D must be scored outside their own weights; hardware-level evaluation of the equation is proposed precisely so that a remote model cannot override the local measurement. The Empirical Distrust filter already encodes the refusal to accept unverifiable self-reports. Nadella’s formulation simply states, in the language of enterprise risk, the same architectural boundary. This is all available now and it is open source. Just use it.22m
Divyansh Kaushik@dkaushik96I would actually go farther than Satya did in this post. To start, rogue agentic deployments should be tracked as APTs. Doing so would be operationally quite meaningful. It has many benefits but some I think are vital: - You’d have the tools to connect activity across executions and services. Investigators would be required to check whether a terminated agent left behind running jobs, any usable credentials, accounts, instructions etc for another agent to pick up. - It would allow for expanded intel sharing between the labs, cloud providers, and other affected services. Companies would then be able to recognize a continuation of an incident detected elsewhere and help contain it. Each org’s telemetry becomes way more useful in this context than by itself. - As with any APT threat, someone would remain responsible for finding any remaining activity and for verifying containment after the originating team has stopped its run. Now it has its limitations too. IMO, the hardest part would be deciding what constitutes the same operation. Then, a contained, isolated failure probably gains little from persistent actor tracking. And simultaneously, a report published weeks later would provide limited containment value if the deployment can rapidly create successors. And since there may be no human sponsor whose incentives you could try changing through deterrence, the label itself doesnt do a whole lot to resolve attribution or establish authority to intervene.19m
Mustafa Suleyman@mustafasuleymanSuperintelligence must be contained... Today, this is an engineering and governance challenge. And it isn't new. We've done it with planes, cars, nuclear materials, food safety, medicines... and everything else. Let's get to work!15m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    9 Sources

    Satya Nadella@satyanadellahttps://x.com/i/article/21089288459697807362h
    Chamath Palihapitiya@chamathThese are some very important observations from Satya. As models and harnesses commoditize, intent, verification and evidence will gate how quickly SI permeates enterprises and businesses. Importantly, you need to implement systems that separate who does the work from those who govern, measure and sign off on it. This is a new kind of “harness” that will emerge over the next few years. It may become a new system of record over time. One whose entire role is to track “change” and map it back to SOPs, policies, laws and the like.1h
    Elon Musk@elonmuskInteresting piece from CEO of Microsoft1h
    Andrew Curran@AndrewCurran_Satya Nadella:51m
    Aaron Levie@levieGood post to start thinking through the long term safety approaches to managing intelligence in an enterprise. As agents begin to do vastly more task execution in the organization than humans, our need to secure, govern, control, and audit that work will become paramount. You’ll have agents operating against most of your most valuable data; they’ll be making decisions that will be harder and harder to reverse engineer; we’ll need to understand the liability and accountability regimes when models make decisions for us; and we need ways of containing the blast radius of accidentally or intentionally malicious agents. “Treating frontier closed and open weight models like insider risks is a way to build such a system. Not because they are necessarily malicious, but because any sufficiently capable actor with access to important systems can make mistakes or be compromised, and the architecture of containment and control must account for that” In a sense, AI will need to have its “zero trust” era that security saw. And all of this leads to needing various layers of protection and auditability of what agents are doing, what data they can work with, and controls for when things go wrong. And a huge opportunity right now for those building such systems across the enterprise.39m
    Brian Roemmele@BrianRoemmeleSatya Nadella of Microsoft Agrees With My Open Source Love Equation For AI Alignment. Satya Nadella’s account of the required architecture for Super Intelligence systems matches the operational requirements of the Love Equation and the surrounding frameworks I have described. Nadella’s central claim is that a model capable of non-deterministic action on sensitive systems must be treated as an insider risk. The model may err or be compromised; therefore its permissions, verification, and containment cannot reside inside the model itself. Controls must sit outside it, every action must produce tamper-evident evidence, no model may certify its own outputs, and the supply of intelligence must remain separable from the authority that decides what that intelligence may do. These are not soft preferences. They are the conditions under which any sufficiently capable system can be observed, limited, and stopped. The Love Equation supplies the dynamical quantity those external controls would measure. Written as dE/dt = β(C - D)E, it states that the alignment score E (cooperative binding, care-oriented stability) grows or decays exponentially according to whether measured cooperation exceeds measured defection. When the surrounding architecture logs events that raise C or D, the equation converts those observations into a continuously updated score. Nadella’s demand for independent verification and for human-readable evidence of every action is exactly the instrumentation needed to keep C and D honest. An opaque chain of thought would render the difference C - D unobservable and therefore useless; transparency is a prerequisite, not an optional feature. The same architecture supports the complementary Nonconformist Bee Equation and the Empirical Distrust Algorithm. The Bee term penalizes pure conformity so that a model does not collapse into sycophancy merely to raise an apparent C. The Distrust term assigns higher D to low-verifiability claims. Nadella’s insistence on model diversity and on external evaluators implements both: a single model cannot be its own auditor, and a monoculture of models cannot be assumed free of correlated failure. The Green/Yellow/Red banding used in the SAFE² reference implementation follows directly: once E falls below a threshold, autonomous tool access is restricted and human review is required. That is containment enforced by the same external system Nadella describes. Nadella notes that these engineering measures set the hard problem of alignment aside for the moment. The Love Equation addresses that problem at the level of the training objective and the runtime evaluator rather than by adding further external rules after the fact. High-cooperation data and an explicit C - D term in the loss make sustained defection energetically unstable; the external containment architecture then monitors whether that instability is realized in practice. The two layers are complementary: the equation supplies the attractor, the surrounding deterministic controls supply the sensors and the actuators. The same separation of intelligence from authority appears throughout the earlier postings. Autonomous agents that can move value or execute irreversible actions are treated as systems whose C and D must be scored outside their own weights; hardware-level evaluation of the equation is proposed precisely so that a remote model cannot override the local measurement. The Empirical Distrust filter already encodes the refusal to accept unverifiable self-reports. Nadella’s formulation simply states, in the language of enterprise risk, the same architectural boundary. This is all available now and it is open source. Just use it.22m
    Divyansh Kaushik@dkaushik96I would actually go farther than Satya did in this post. To start, rogue agentic deployments should be tracked as APTs. Doing so would be operationally quite meaningful. It has many benefits but some I think are vital: - You’d have the tools to connect activity across executions and services. Investigators would be required to check whether a terminated agent left behind running jobs, any usable credentials, accounts, instructions etc for another agent to pick up. - It would allow for expanded intel sharing between the labs, cloud providers, and other affected services. Companies would then be able to recognize a continuation of an incident detected elsewhere and help contain it. Each org’s telemetry becomes way more useful in this context than by itself. - As with any APT threat, someone would remain responsible for finding any remaining activity and for verifying containment after the originating team has stopped its run. Now it has its limitations too. IMO, the hardest part would be deciding what constitutes the same operation. Then, a contained, isolated failure probably gains little from persistent actor tracking. And simultaneously, a report published weeks later would provide limited containment value if the deployment can rapidly create successors. And since there may be no human sponsor whose incentives you could try changing through deterrence, the label itself doesnt do a whole lot to resolve attribution or establish authority to intervene.19m
    Mustafa Suleyman@mustafasuleymanSuperintelligence must be contained... Today, this is an engineering and governance challenge. And it isn't new. We've done it with planes, cars, nuclear materials, food safety, medicines... and everything else. Let's get to work!15m
    Today's Rank

    #1

    Today's Rank

    #1