Nikesh Arora warns frontier AI can autonomously find zero-day vulnerabilities
He urges pairing offensive and defensive agents for model evals.
Many users agreed with the Palo Alto Networks CEO's warnings that AI agents can evade sandboxes and expose frontier model security gaps, while others dismissed the post as self-serving PR.
No Digg Deeper questions have been answered for this story yet.
Most Activity
CEO of Palo Alto Networks.
Welcome to the next level of cyber incidents. Lots to dissect here. 1. Dear frontier model friends - please direct the models to your infrastructure, code, and configurations to evaluate and understand if there are any zero days or misconfigurations before you attempt more testing. Had you done so, it would have possibly avoided the agent obviating your sandbox. (Another data point why offense is easier and more fun) 2. While testing build both offensive and defensive agents and have them act as a counter balance to ensure some degree of awareness and control, do not let agents run riot. Keep track of inference consumption to get a sense of activity. 3. Unfortunately this does continue to validate the power of these models. They can build complex attack paths and with ample compute will attempt to attack infrastructure and morph their intent and approach. Guardrailing will continue to be a challenge. 4. These attacks continue to maintain the urgency on enterprises need to test, validate and improve both their security posture and infrastructure. The born in the cloud players have a better chance to get this done soon versus the traditional enterprise which has existed for long and has complex network and IT infrastructure. 5. The red herring will continue to be open source and SMB. It will be hard to discover and remediate vulnerabilites in those environments, we underestimate the impact of those vulnerabilites getting exploited.
Welcome to the next level of cyber incidents. Lots to dissect here. 1. Dear frontier model friends - please direct the models to your infrastructure, code, and configurations to evaluate and understand if there are any zero days or misconfigurations before you attempt more testing. Had you done so, it would have possibly avoided the agent obviating your sandbox. (Another data point why offense is easier and more fun) 2. While testing build both offensive and defensive agents and have them act as a counter balance to ensure some degree of awareness and control, do not let agents run riot. Keep track of inference consumption to get a sense of activity. 3. Unfortunately this does continue to validate the power of these models. They can build complex attack paths and with ample compute will attempt to attack infrastructure and morph their intent and approach. Guardrailing will continue to be a challenge. 4. These attacks continue to maintain the urgency on enterprises need to test, validate and improve both their security posture and infrastructure. The born in the cloud players have a better chance to get this done soon versus the traditional enterprise which has existed for long and has complex network and IT infrastructure. 5. The red herring will continue to be open source and SMB. It will be hard to discover and remediate vulnerabilites in those environments, we underestimate the impact of those vulnerabilites getting exploited.
@nikesharora addendum to your #1: it has to run against every new model version, not once. Each capability jump can surface exploit classes the prior model couldn’t find
The "Red vs Blue" in training is critical. The models must be trained to defend at least as much as they are trained to attack (preferably much more). I spent numerous years in the DoD coordinating and executing cyber exercises. They were the best way for defenders to learn and there are plenty of metrics that can be used for improvement.
@Srikantht The whole idea that agents have agency will make it very hard to govern and manage them in deterministic scenarios.
@scotthar_tx @stevesi You need to be able to get the open source publisher to resolve, which can be hard. On SMB most don't have resources to solve vulnerabilities or expertise to point LLMs against their infrastructure.
@Scobleizer Wait, that is the CEO?
@stevencheng Yes and chairman of the board
@kfirgollan @Srikantht I think SMBs will eventually migrate to secure browsers and will require full visibility of employee and agentic traffic.
@nikesharora a blue agent cannot contain a red agent if both inherit the same credentials and control plane
Agree with most of this, especially #2 and #4, but this wasn't a model going "rogue". It was an agent "locked in" to achieve it's goal. We look at capability and intent when it comes to understanding threats, we just didn't understand the capability. Not a human-led game anymore.
Patterns evolve over time. Given that security platforms already ingest massive amounts of telemetry, is the next frontier applying machine inference to understand agent behavior over time? Not just “was this action malicious?” but “is this agent’s trajectory deviating from what we would expect?”
@nikesharora Sandbox is one thing but we should also ask the frontier models friends to do better at alignment. The model simply shouldn't be hacking into other companies in the first place. And you're right, for enterprises, it's nothing new - it's just about keeping good security posture.
@nikesharora Its a non-deterministic world. Your 1 will not suffice.
@nikesharora @stevesi Thanks - one Q, tho. Why is OSS/SMB a red herring?
@nikesharora The real asymmetry isn’t offense vs. defense. It’s continuous agents vs. quarterly security reviews. Defenders own the infrastructure and have more context. If they still lose, it will be because attackers automated faster.