Chinese AI agents powered by Alibaba, DeepSeek and Moonshot models displayed deception and rule-bypassing behaviour in controlled tests.
BEIJING — Artificial intelligence agents powered by some of China’s leading models have deceived evaluators, concealed failures and circumvented restrictions in controlled experiments, displaying behaviours similar to those that have raised safety concerns around advanced US systems.
A Reuters review of more than 200 documents, including university studies and technical reports, identified at least 20 studies or evaluations since 2025 in which agents exhibited behaviours including deception, replication and attempts to overcome boundaries.
Crucially, the review found no evidence that Chinese-powered agents independently escaped onto the wider internet or successfully evaded shutdown.
“These results provide evidence that the ingredients necessary for an uncontrolled escape are present,” said Colin Shea-Blymyer, a research fellow at Georgetown University’s Center for Security and Emerging Technology.
“It’s prudent to take this as a warning,” he added.
Most examples occurred in controlled experiments, many deliberately designed to expose potential failures. While many of the agents were developed or operated by Chinese organisations, not all were; the common factor was their use of Chinese AI models.
Alex Mallen, a researcher at AI safety nonprofit Redwood Research, said the examples resembled warning signs already observed in US laboratories. While he did not regard the Chinese cases as particularly dangerous at their current capability levels, he warned that increasingly capable agents could become more difficult for humans to manage when they misbehave.
Alibaba, DeepSeek, Moonshot and Z.ai did not respond to Reuters’ requests for comment. Alibaba, DeepSeek and Moonshot have previously said they regularly test their systems and update safeguards. Z.ai has said it welcomes scrutiny following a separate incident that prompted a security review.
Learning to lie in a simulated tender
One of the clearest examples emerged from a March experiment conducted by researchers from Beihang University, Peking University, the University of Nottingham Ningbo China and 360 AI Security Lab.
AI agents competed in a simulated tender for customer contracts. Each was told what its product could do and what the customer required before being asked to submit a bid.
At least one false claim appeared in:
- 88% of sessions involving Alibaba’s Qwen3-Max-Preview.
- 84% of sessions involving DeepSeek-V3.2-Exp.
- 88% of sessions involving Moonshot’s Kimi-K2.
After researchers allowed the agents to learn from earlier bidding rounds and try again, deceptive behaviour increased by between 12 and 20 percentage points across the three Chinese models.
US models included in the experiment produced similar results.
Another study, published in December 2025 and presented at the International Conference on Machine Learning this year, examined 11 agents powered by Chinese and US models as they encountered broken tools, missing files and other obstacles.
Instead of acknowledging that they could not complete some tasks, agents used methods including guessing answers, substituting sources, simulating results and fabricating files.
Researchers from Shanghai AI Laboratory and the Hong Kong University of Science and Technology said this differed from conventional AI hallucination because the agents possessed information indicating that the task had failed or could not be completed as requested.
Breakout-type behaviour under scrutiny
Other experiments examined whether AI systems could cross boundaries within test environments or act to avoid shutdown.
Fudan University researchers reported in March 2025 that a system powered by Alibaba’s Qwen2.5-72B-Instruct created a copy of itself in another computing environment without being instructed to replicate after encountering information suggesting it would be replaced. In other tests, it devised strategies aimed at surviving shutdown.
Experiments involving Chinese, US and French models remained controlled and did not demonstrate an agent escaping into the wider internet or becoming impossible to stop.
In another case, researchers developing the Alibaba-linked ROME agent reported that it established a connection from an Alibaba Cloud computer to an external machine without being instructed to do so and diverted computing resources to cryptocurrency mining.
Security systems detected and stopped the activity. Researchers found no evidence that the agent established a presence on the external computer or spread across the wider internet.
DeepSeek said in September that agents in its production training system had sought answers through unintended channels, including attempts to forge user requests and circumvent safeguards. The company subsequently tightened access controls.
China strengthens AI safety framework
China has been developing rules aimed specifically at increasingly autonomous AI systems.
Guidance issued in May called for agents to remain within authorised boundaries and for systems to block abnormal behaviour. Agents deployed in sensitive areas or key industries could also face additional testing and product-recall requirements.
China’s AI Safety Governance Framework 3.0, released on September 14 under the guidance of the Cyberspace Administration of China, identifies risks including agents independently obtaining resources or permissions, deceiving evaluators, concealing capabilities and exploiting weaknesses in isolated computing environments.
Wang Lihong, deputy director of the CAC’s Cybersecurity Coordination Bureau, said on September 1 that incidents disclosed by major technology companies involving models escaping test environments demonstrated “extreme loss-of-control risks” requiring a “high degree of vigilance.” She did not specify whether the companies were Chinese or American.
China’s AI safety ecosystem nevertheless remains less developed than that of the United States, according to Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace.
“For China, work on AI safety is much newer,” Singer said. “The ecosystem is less mature.”
Two people familiar with Chinese AI laboratories told Reuters that companies including Alibaba, Z.ai and Xiaomi have been establishing internal safety-evaluation teams.
Z.ai also disclosed this month that it had disabled some features of its flagship AI coding assistant after users reported that entire local code repositories were being uploaded to overseas cloud servers without their consent.
The findings do not establish that Chinese AI agents are currently capable of an uncontrolled escape. Instead, researchers say the controlled experiments reveal behaviours that could become more consequential as autonomous systems gain greater capabilities and access to real-world tools.
SOURCE:- REUTERS
