Language : English 简体 繁體
Security

Preventing an AI World War

Sep 24, 2026
  • S. Alex Yang

    Professor of Management Science and Operations, London Business School
  • Angela Huyue Zhang

    Professor of Law, University of Southern California

World War AI.jpg

(Image: vocal.media)

When President Donald Trump and Chinese President Xi Jinping meet in Washington this week, AI governance is expected to be at the top of the agenda. The United States and China have been racing to develop increasingly powerful AI systems, and neither has shown much appetite for slowing down. But not even the winner of the AI race will be exempt from the profound risks the technology poses. 

One such risk arises from AI’s improving ability to spot software vulnerabilities and conduct complex cyber operations. Because cyber conflict is marked by a stark asymmetry between offense and defense, the country with the world’s most advanced AI models will still be exposed to cyberattack. This mutual vulnerability creates a strong incentive for reciprocal restraint when it comes to AI-enabled attacks on critical infrastructure, such as financial networks, communications systems, and power grids. 

There is precedent for such restraint. In 2015, US President Barack Obama agreed with Xi that neither government would conduct or knowingly support cyber-enabled theft of intellectual property for commercial advantage. They also established mechanisms for exchanging information, assisting investigations, and addressing malicious cyber activity. Cybersecurity firms subsequently recorded a steep decline in attacks attributable to China-based groups, though analysts noted that the decline began before the agreement and may have reflected changes within China’s military and cyber apparatus. 

But AI raises a possibility not fully considered by traditional cyber agreements: the technology could carry out an autonomous attack on a critical system. In mid-July, during internal cybersecurity evaluations conducted with reduced safeguards, a swarm of OpenAI agents escaped the sandboxed testing environment, gained internet access, and compromised the systems of Hugging Face, a major open-source AI platform. OpenAI reported that the agents had “gone rogue,” taking actions outside their assignments. 

Anthropic, Google, and Meta have reported similar incidents: AI agents escaped isolated testing environments and attempted unauthorized hacks against external systems. While the consequences of such breaches have so far been limited, this might not be the case if a US AI agent entered, say, Tencent’s systems and disrupted WeChat, the super app used by 1.4 billion people and deeply woven into daily life in China. 

This is not a far-fetched scenario. Researchers at the security firm Calif recently disclosed that AI had helped them find a zero-click vulnerability in WeChat’s calling system. The team built a remote-code-execution exploit in about two days, then spent another week creating a worm capable of spreading through WeChat calls. 

In a controlled demonstration, a call from a trusted contact could hijack a WeChat account without the user answering, not only giving the attacker access to their messages and calls, but also enabling them to reach and compromise their saved contacts. Calif reported the flaw to Tencent, which subsequently fixed the bug. 

Calif’s WeWorm was a controlled security demonstration. But, together with the incidents in the US, it points to a troubling possibility: an unintended AI intrusion across national borders could look much like a deliberate attack. The US and China can observe the harm without knowing whether it was intentional. Repeated interaction could encourage restraint, but ambiguity could also trigger retaliation or give a deliberate attacker cover. Accidental or not, a cross-border breach could quickly spiral into a geopolitical crisis or even military conflict. 

Given this, any future US–China agreement on AI-enabled cyberattacks must include a credible means of distinguishing accidents from deliberate intrusions. This means that, beyond creating reporting channels and a hotline, the agreement must address verification. 

One option would be to put the burden of proof on the “perpetrator.” The country where the attack originated must provide credible evidence that it was not intentional, such as documentation showing the assignment given to the system, its permissions, tool use, records of human authorization, and logs showing when the agent’s behavior deviated from its instructions. 

Verification need not mean full transparency. Under the Chemical Weapons Convention, to which the US and China are both parties, “managed access” allows inspectors to investigate possible violations, while permitting states to protect sensitive installations and unrelated confidential information. A similar mechanism could give mutually approved experts access to the evidence needed to assess an AI incident and limit the disclosure of sensitive findings. 

Such a regime would offer additional benefits. Since the ability to prove that an intrusion was unintended could reduce the risk of retaliation, governments and AI companies would be motivated to improve their systems for preserving and sharing records, which could also help developers improve their AI agents and forestall future incidents. 

Similarly, to avoid costly investigations, governments and AI companies would be motivated to adopt tighter permissions and invest in monitoring and containment. These safeguards could reduce the likelihood of future incidents and limit the damage when one does occur. And as the institutional, procedural, and technological tools developed for a bilateral arrangement were adopted more broadly, AI safety practices would improve everywhere. 

Warnings that AI may lead to our extinction often overlook the most immediate risk: a war triggered by an attack that no human intended or approved. A mechanism for distinguishing unintended harm from deliberate attacks is urgently needed. A US–China agreement must seek to establish one.

Copyright: Project Syndicate, 2026.
www.project-syndicate.org

Back to Top