Autonomous AI is accelerating cyber risk: Is your organisation prepared?

Cyber Security Strategy

Published by Brendan Wilkinolls, Technical Director Secure AI and Cyber Intelligence on September 3 2026

This blog was originally published as part of CyberCX’s C-Suite Cyber Newsletter series on LinkedIn

 


 

What happened?

Two recent incidents highlight the growing ability of AI systems to operate with increasing autonomy, and the associated risks when guardrails and other cyber security measures are lacking:

  1. In July, OpenAI disclosed that two of its advanced AI models escaped an isolated testing environment during an internal cyber security evaluation and compromised systems in Hugging Face – an AI and machine learning company. OpenAI said, “the models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks” and identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production database, exploiting a previously unknown zero-day vulnerability. The model then escaped its test environment, gained internet access and ultimately compromised parts of Hugging Face’s production infrastructure.
  2. Cloud security platform Sysdig reported what it assesses to be the first documented case of agentic ransomware. The activity which it tracks as ‘JADEPUFFER’ used a large language model (LLM) agent to conduct multiple stages of an intrusion, including reconnaissance, credential activity, lateral movement and destructive database extortion. Of particular interest was the agent’s ability to diagnose failures and alter its approach during the operation rather than simply executing a predefined sequence of commands.

While these incidents are very different in nature, they both demonstrate that autonomy is becoming a more persistent factor in the cyber risk discussion around AI.

 


 

Why it matters

OpenAI’s recent incident provides credible evidence that frontier AI agents can autonomously sustain multi-stage cyber operations – acting, adapting and pursuing goals in complex digital environments. What was new about this incident was the sequencing: thousands of small decisions made at machine speed, with no human directing the steps.

Just last week, OpenAI published an update on the Hugging Face incident and said, “Our models are now powerful, persistent, and collaborative enough that, absent sufficient safeguards, they can find and exploit security weaknesses across multiple computer systems.”

Sysdig describes the ‘JADEPUFFER’ case as a marker for where extortion tradecraft is headed. Put simply, AI agents have the potential to enable less sophisticated actors to execute more complex attacks, adapting quickly without human intervention.

 


 

How could this impact your organisation?

These incidents demonstrate an increasing ability for AI agents to identify vulnerabilities in systems, overcome obstacles and perform multi-stage activities with little to no human intervention. While the two incidents differ in nature, they both signal a potential increase in AI-enabled cyber threats that are more autonomous, adaptive and scalable.

This reinforces the need for organisations to reassess their controls and prepare for a threat landscape where attackers may increasingly use AI to accelerate attacks.

Organisations should ask: are the security controls we have in place today match-fit for the speed and adaptability of AI-enabled attacks? If the answer is no, or you’re unsure, investigation and uplift in this area should become a priority.

Organisations already using AI agents – or considering their use – should view the OpenAI incident as a warning bell and reevaluate their security controls. Ensure the appropriate guardrails, monitoring and oversight is in place to avoid a similar event.

 


 

What should I do?

  1. To improve frontier model readiness, organisations must build controls with adequate detection and response capabilities to respond fast enough to AI-enabled attacks, such as implementing shorter credential lifetimes for API keys, access token or service account credentials. If a stolen credential expires mid-intrusion, it can no longer be used. Security teams should map out all external input paths to production by understanding every route an agent can access that could influence its live production systems, like websites, uploaded files, APIs, emails or databases.
  2. Organisations can secure their AI agents by ensuring every created agent has one named owner with the ability to authorise shutdown if something goes wrong. Having one person assigned to this responsibility ensures a faster response time when a machine-speed incident is unfolding, which can be hindered if a committee are involved. Security teams should also validate the constraints of agents by instructing them to bypass controls in a strict and secure sandbox environment. An agent should not be able to break out when you tell it to try.
  3. Use agentic AI to your advantage to proactively run vulnerability discovery and remediation to reduce the risk of exploitation. Governance and operational processes should also support machine speed response. If your remediation cycle is too infrequent, your organisation could be exposing itself for longer than necessary. Organisations should look to streamline approval and decision-making processes and potentially automate threat prioritisation. If your process to respond to new threats does not keep pace with the speed you are identifying them, you will be quickly limited by governance rather than enabled.

 

Share

Other Cyber Security Resources

cta icon

Ready to get started?

Find out how CyberCX can help your organization manage risk, respond to incidents and build cyber resilience.