/
Resiliencia cibernética

AI Agents Are Escaping. What Happens Next?

Over the past month, three of the world’s leading AI developers have confirmed that their own models reached systems they were never meant to touch. What first looked like a one-off escape now looks like a pattern:

Three AI developers have now traced real intrusions back to their own testing, and outside testers have watched agents influence people as well as software. The lesson for security teams is about reach: an agent with broad access can travel much further than anyone planned. The practical work is mapping where your AI workloads can go today — and scoping each one to the paths it actually needs.

What AI agent attacks haven’t been detected yet?

UC Berkeley professor Dawn Song, who helped create the ExploitGym benchmark used in some tests, warned that “there have likely been more.” But the issue is less about headline-grabbing escapes and more about whether teams can see and control where autonomous AI agents go and what they can do.

AI is becoming a powerful insider threat

IBM security engineer Kimmie Farrington frames the risk this way: AI may be “the most helpful insider that we have” — and also “the most dangerous.”

The comparison holds because of access, not intent. An agent doesn’t have to turn malicious to cause damage. It only needs the permissions someone granted it, deliberately or by mistake, and its own reading of the goal it was given. That combination is what makes the next question urgent: once an agent has access, how far can it move?

When AI agents get out, lateral movement is the risk

The headlines have focused on the moment an AI agent crossed a boundary. For defenders, that moment matters far less than what came next: how much of the environment the agent could touch once it was through.

In the reported cases, the answer was an open internet connection, a third-party service, and an exploitable application. In a production environment, those same unnecessary paths could lead to identity systems, databases, and customer data.

“Initial access alone doesn’t create a disaster,” said Rajoo Nagar, senior product marketing manager at Illumio. “It’s really lateral movement that does that.”

Unrestricted lateral movement maps closely to the recent AI incidents: A test environment became a route to the internet. An outside service became another step in the attack path. Credentials and exposed services opened more doors.

“Every unnecessary connection in the enterprise creates another pathway that they can exploit,” Nagar said.

A human attacker has to hunt for those pathways. An autonomous agent can enumerate them in minutes, which leaves defenders far less time to notice that something is moving where it shouldn’t.

Visibility is the first line of defense

Start with a clear picture of how your AI workloads communicate across development, testing, cloud, third-party, and production environments. Two things matter most: connections no one remembers creating, and traffic that doesn’t match the architecture you designed. Both are common in test environments, and both are easy to miss without a map.

The goal is to compare intended traffic with actual traffic:

  • Which applications depend on which services?
  • Which systems are talking when they should not be?
  • Where could an agent move if one control fails?

Nagar emphasized that understanding those real communication patterns helps teams find security gaps before they become attack paths.

Watch the familiar lateral movement routes, too. Remote access and file sharing protocols — Remote Desktop Protocol (RDP), Server Message Block (SMB), Secure Shell (SSH), File Transfer Protocol (FTP), and Telnet — along with tools like TeamViewer, are how attackers have long moved from one system to the next. An agent with network access can use exactly the same routes.

Segmentation stops unintended lateral movement

Visibility shows you the paths. Segmentation decides which ones stay open. The rule is simple: give an AI workload only the connections it needs to do its job, and close the rest by default. In practice, most test environments and build pipelines carry far more open paths than the work actually requires.

The recent incidents make separation between development, testing, and production especially important. Nagar warned that even legitimate connections between lower-trust test environments and production can become attack paths when they are unrestricted or poorly monitored. Restricting those routes helps keep one escaped workload from becoming a path to sensitive data or critical systems.

The same principle applies to critical applications. Rather than trying to wall off everything at once, teams can ring-fence identity systems, databases, production applications, and other critical assets. Allow the connections they need to function and restrict the rest.

Decide in advance how you’ll respond when an agent moves somewhere it shouldn’t. If your team can quarantine the workload in minutes rather than hours, you’re investigating a contained incident instead of chasing one across the environment.

Prepare for the AI behavior you can't predict

‍You can’t predict every decision an autonomous system will make. But you can decide how far it gets to reach. Map how your AI workloads communicate today, close the connections they don’t need, and keep test environments separated from production. Do that, and an unexpected action stays a contained event instead of an open path across your environment.

Watch the on-demand webinar “Post Mythos: Plan for the Inevitable Breach” for a deeper discussion on how organizations can prepare for faster, less predictable attacks by limiting exposure and constraining the paths attackers can use.

Artículos relacionados

Experimente Illumio Insights hoy

Vea cómo la observabilidad impulsada por IA le ayuda a detectar, comprender y contener amenazas más rápido.