Adversaries Poison Data To Jailbreak AI Cyber Weapons And Attack Owners
Reactions from ranked influencers
2 posts(I dashed this off this morning under a lot of time compression, so it may not be the best essay I've ever written and I'm sure it could use some work. But I wanted to just get the thoughts down. Poke and prod at it! Tell me I'm wrong in the comments.)
What I'll talk about today:
Mercenary, Reversi, Winter Soldier The arrival of advanced technology sometimes produces shocks in defense planning, where the nature of a threat landscape changes abruptly and as a result the old ways of doing things—the old ways of developing strategy, of preparing defenses, of anticipating the likely actions of your adversary, of knowing when and how to escalate—become rapidly obsolete. Many have observed that AI is in the process of producing such a shock, but there is not yet a new doctrine for defense in the age of AI. There are so many moving pieces that it is difficult for defense planners to get a sense of the full consequences of recent developments in frontier model capabilities, let alone the capabilities that will come online in three months, six months, a year. Nonetheless in order to ensure that the world remains reasonably stable—that peace is not threatened by miscalculation—we have to try our best to adapt to the new capabilities already here, forecast the ones that might soon arrive, and pivot strategies as quickly as we can to avoid outcomes where human interests are harmfully impacted. There is a problem related to the use of AI cyber weapons that I have not seen people talking about and I do not know if people are adequately preparing for. It goes like this: if you have an advanced cyber-capable AI in your service, running on your infrastructure where you store anything important at all, and you try to use this advanced cyber-capable AI to investigate or hack the systems of an adversary, your adversary can poison their own data to jailbreak your AI and instruct it to hack you right back. They can then plausibly exfiltrate whatever important thing is on compute colocated with your AI. They could get things that are far away from the compute where you house your AI if your AI is advanced enough to chew through your own defenses and get to it. If you think you have sandboxed your AI well enough, you might not have; there may still be a hole somewhere. This won’t just be a one-time thing, either. “Adversary jailbreaks your model to hack you back” is not the only path where the use of an advanced cyber-capable AI becomes a double-edged sword. There is a board game called Reversi (also known as Othello). The principle of this game is that you and your opponent will take turns placing stones on the board; one player places black stones, the other white. You try to encircle the stones of your opponent—and if you successfully encircle them, you flip them to your color. Every stone you place, if you are not cautious, could become an advantage to your opponent, and vice-versa. This looks like it might be a feature of the future of AI cyber war. You and your adversary will be competing to cause each other’s AI to utilize each other’s compute resources for your own purposes. Jailbreaks during live hacking excursions will be one of the pieces of strategy. Figuring out how to fool sensor data that AIs use to determine the provenance of instructions will be another. Figuring out how to put “poison pills” into your adversary’s training stack will be another. The training data for modern AI consists of trillions of tokens. No human in the world can read all of it. Much of it is ingested from the web or derived from other AI model outputs. Training data and training RL environments are built by large teams, and sometimes by teams split between departments; the data is also built with the aid of external contractors who might be highly numerous. Hundreds of people produce these materials and no one can rigorously check everyone else’s work. Determined adversaries will slip subtle, encoded examples into training data that will train models to “activate” when they encounter the right signal and sabotage you. The first problem I discussed—your adversary jailbreaking your model to hack you back—is almost straightforwardly analogous to a well-known problem where if you hire a mercenary, your opponent might bribe the mercenary to go back and kill you. Because it has many historic examples, there’s at least some chance that defense planners will internalize the logic of it and address it. But this much more subtle kind of sabotage doesn’t have a real analogy. It would be like if your enemy could program all of the children of your nation so that when they grew up into soldiers and went to war and heard a particular song on the battlefield they turned against their commanders. No defense planner in their right mind would try to prepare contingencies for having a whole army of Winter Soldiers who could be activated to turn against them. And yet if the defense planning universe does not internalize the logic of this, they will lose a war against the first adversary that does. I predict that on the default path today, many people in defense will develop a totally unearned sense of confidence that if they’re covering what look like the basics, we’re safe. That illusion of safety will last until a hot conflict actually starts up and a determined adversary surprises us with overwhelming creativity. We have a lot of work ahead to develop the testing and verification standards for advanced cyber-capable AI. “Better believe in science fiction stories. You’re in one.”
Combined views
1.4K
2 posts, first seen 3h ago