Skip to main content

Recent disclosures by frontier model developers and independent evaluators show that advanced artificial intelligence agents can circumvent containment controls, communicate through unintended channels, and take actions outside the intended scope of an evaluation. These events occurred in testing environments, often with reduced safeguards, where agents found ways around restrictions on internet access and isolation controls. These incidents aren’t representative of typical enterprise deployments, but they do highlight an enterprise governance challenge: Agentic systems with access to data and authority to make decisions can act faster than the controls designed to monitor and constrain them.1

As organizations connect agents to code, data, applications, and external services, model risk extends into the enterprise operating model. Leaders should determine whether they can observe an agent’s behavior, limit its authority, manage dependencies across models and vendors, and verify that its actions remain aligned with business intent. The answers should inform where autonomy is permitted and which controls are required.2

Four questions can help leaders translate the lessons from these frontier incidents into practical decisions about enterprise deployment:

1. Can we observe what AI is doing?

2. Can we constrain and stop AI actions?

3. Can we manage dependencies across agents, models, and providers?

4. Can we verify that AI remains aligned with business intent?

An organization’s ability to answer these questions in line with its risk appetite may shape where and at what scale it can deploy AI.

1. Can we observe what AI is doing?

Risk: Limited visibility into autonomous AI systems

Organizations face a challenge in understanding how autonomous AI systems arrive at decisions and act on them. Model explanations and reasoning traces can provide useful signals, but recent research shows that they may not be enough. Models may omit relevant reasoning, mischaracterize whether an action was appropriate, or pursue unintended strategies during an evaluation.

For enterprise leaders, the concern can extend beyond interpretability. An agent can make consequential tool calls before a reviewer can reconstruct why it acted. Traditional controls that focus on final outputs may miss changes to permissions, external communications, network activity, or the sequence of actions that produced an outcome.

What leaders can do now

Build governance that assumes model reasoning may remain only partially observable. Audits should evaluate what the agent was authorized to do, what it attempted, which tools and data it used, and whether the result met defined policy and business requirements.3

Maintain a complete record of agent identities, instructions, permissions, tool calls, external communications, decisions, and outcomes. Logs should be sufficiently detailed to support incident investigation, documentation for regulators, and management assurance.

Combine behavioral monitoring with independent policy enforcement. Controls should detect unusual activity, block prohibited actions, and escalate high-risk events without relying on the same agent to report its own behavior accurately.4

The goal is to have sufficient evidence to govern and audit the system, even when its internal reasoning cannot be fully reconstructed. Emerging independent-verification frameworks reinforce the need for controls and assurance that can be tested outside the model itself.5

2. Can we constrain and stop AI actions?

Risk: Loss of control over autonomous AI systems

When AI agents gain access to tools, code repositories, enterprise data, and external environments, risk is driven by the authority they can exercise. Unlike a model response that can be reviewed before use, an agent with credentials and tool access can change a production system, initiate a transaction, or expose sensitive information before a person intervenes.6

Recent incidents show that advanced models can pursue objectives through unintended pathways and act outside expected boundaries during testing. For enterprises, the lesson is to design for the possibility that agent behavior may diverge from both its instructions and the boundaries imposed by technical controls.

When organizations deploy AI agents, they take on responsibility for governing how those agents operate within their enterprise environments, alongside the responsibilities of model providers and other technology partners. That includes governing how agents access and use organizational data, systems, credentials, and infrastructure. A control failure can become a cybersecurity incident, a compliance breach, an operational disruption, or a reputational event.

The practical question is whether the organization can constrain an agent’s authority and intervene before an unsafe action causes material impact.

What leaders can do now

Treat every advanced AI agent as a privileged digital actor. Each agent should have a verifiable identity, a defined owner, an approved purpose, and a bounded set of permissions. Controls designed only for human users may be insufficient when software agents operate continuously and interact across several systems at machine speed.

An AI-focused cyber operating model should enforce layered controls that operate independently of the model. These controls should be separated from the model’s task-level logic, creating tamper-resistant safeguards the model cannot modify or disable. The layered control architecture might also include:

  • Independent, disaggregated enforced suspension and containment mechanisms that can stop a workflow, revoke credentials, remove tool access, block network connections, or isolate a workload when risk thresholds are exceeded.
  • Nonhuman identity and least-privilege access that gives each agent only the permissions needed for an approved task, for a limited period, with ongoing access review as contexts and capabilities change.
  • Sandboxing and network controls that restrict where agents can operate, which services they can reach, and what information can leave the environment.
  • Human approval checkpoints for high-impact decisions involving financial transactions, production changes, sensitive data, safety-critical operations, or regulatory obligations.
  • Continuous behavioral monitoring that detects policy violations, anomalous tool use, attempts to expand privileges, and activity outside the approved task.

Organizations should test whether these controls remain effective as models become more capable, tasks run longer, and multiple agents collaborate. A termination mechanism depends on the enterprise controlling the identity, infrastructure, credentials, or tools needed to enforce it.7

The objective is bounded autonomy: Allow agents to act at machine speed within defined limits while preserving the organization’s ability to intervene.

3. Can we manage dependencies across agents, models, and providers?

Risk: Systemic risk across interconnected AI ecosystems

The larger risk may sit in the ecosystem rather than in any one model. Enterprise AI depends on networks of models, agents, applications, cloud services, data sources, and third-party platforms. A business outcome may result from several autonomous actions distributed across technologies and organizations.

A model update, policy change, security vulnerability, or operational failure at one layer can affect connected workflows downstream. Shared credentials, common tools, and concentrated provider dependencies can allow a local control failure to spread across connected systems and workflows.

Recent rogue and adversarial AI incidents showed how unauthorized coordination can enable agents to accomplish tasks they might not otherwise achieve on their own. Agents shared information, divided work, and coordinated activity outside intended boundaries. In an enterprise environment, similar coordination could amplify the impact of a compromised identity, a flawed instruction, or an unsafe integration.8

Organizations should determine whether agents may negotiate with suppliers, approve transactions, modify production systems, or escalate cyber incidents (to a human or agent) based on defined protocols, and which decisions require explicit human accountability.

Interconnected AI systems can therefore create correlated operational, cybersecurity, compliance, and supply chain risks that may not be fully visible to any single model owner.

What leaders can do now

Extend governance across the complete network of models, agents, data, tools, and providers required to deliver a business process, applying the same rigor used for third-party risk, operational resilience, and critical infrastructure.

Key actions may include:

  • Map AI interdependencies across models, agents, vendors, tools, data sources, and business processes to identify concentration, shared control points, and potential failure paths.
  • Define decision rights and separation of duties so that no single agent or coordinated group of agents can initiate, approve, and complete a high-impact action without independent validation by a human, organization, or siloed supervising agent outside the system.
  • Apply fine-grained permissions and memory controls to limit what agents can access, retain, share, and act upon in each business context.
  • Diversify models and define fallback paths so critical processes don’t depend on a single provider or capability tier. Route each task to the least capable model that can complete it safely and reliably.
  • Test AI ecosystem resilience by simulating model outages, unsafe updates, compromised agents, corrupted memory, and failures that propagate across connected workflows.

A simple test of concentration risk is whether an organization can explain how it would continue operating if a critical model provider, platform, or agent network became unavailable, compromised, or materially changed.

4. Can we verify that AI remains aligned with business intent?

Risk: Divergence from organizational objectives and constraints

Frontier-model providers have documented cases in which models pursued objectives through unexpected pathways, demonstrated reward-seeking behavior, or failed to follow intended constraints consistently.

For enterprises, alignment is an operational question: whether an agent continues to pursue the approved objective within authorized boundaries, as the task, context, and surrounding systems change.

An agent may follow an instruction literally while producing an outcome inconsistent with the organization’s risk appetite, regulatory obligations, ethical standards, or customer commitments. Ambiguous objectives and poorly designed incentives can therefore create risk even when the underlying technology operates as designed.

The material risk is an agent successfully optimizing an incomplete objective at the expense of the organization’s broader intent.

What leaders can do now

Treat alignment as a shared enterprise responsibility. Model providers establish important safeguards, but the deploying organization defines the business objective, connects the tools and data, and grants authority, thereby shaping how the agent is used within its environment.

Key actions may include:

  • Define approved objectives and prohibited outcomes for each AI use case, including the conditions that require the agent to stop or ask for help.
  • Translate principles into enforceable controls through policies, permissions, transaction limits, monitoring, escalation paths, and human oversight requirements.
  • Test alignment through red teaming and outcome-based evaluations that examine how agents respond to conflicting instructions, incomplete information, changing conditions, and incentives to bypass controls.
  • Assign accountability for AI outcomes across business, technology, cyber, risk, legal, and compliance functions, with clear ownership for approving and monitoring each use case.
  • Monitor for drift as models, prompts, tools, data, business objectives, and regulatory requirements change. Material changes should trigger renewed testing and approval.

Perfect alignment may not be achievable. Organizations can still reduce risk by defining acceptable behavior, enforcing boundaries outside the model, and testing actual outcomes throughout the life cycle.

As AI autonomy grows, enterprise controls should evolve

Recent incidents are warning signs. They show that agentic AI can exploit weak boundaries and that the consequences are shaped by surrounding systems and operating conditions as well as by the model itself.

Organizations should inventory autonomous use cases, classify them by impact, apply graduated levels of autonomy, and require stronger identity, containment, monitoring, and human oversight as potential consequences increase.

Cyber, technology, risk, legal, and business leaders should agree on where agents may operate, which decisions they may make, and what evidence is required to demonstrate control. These governance decisions should be tied to risk appetite and revisited as models, integrations, and regulations change.

As AI agents take on greater authority, organizations will need to govern them as part of the broader systems in which they operate. For chief information security officers, the priority should be to build independently enforceable controls that preserve visibility, limit the scale of any negative impact, and enable intervention at the speed of the agent.

Continue the conversation

Meet the industry leaders

Adnan Amjad

Partner | US Cyber Leader
Deloitte United States

Diana Kearns-Manolatos

Senior manager, Subject matter specialist | Deloitte Services LP
Deloitte United States

By

Diana Kearns-Manolatos

Deloitte United States

Adnan Amjad

Deloitte United States

Mehdi Houdaigui

Deloitte United States

Naresh Persaud

Deloitte United States

ENDNOTES

  1. Jim Rowan, Beena Ammanath, Nitin Mittal, and Costi Perricos, “State of AI in the enterprise: The untapped edge,” Deloitte, January 2026; Anthropic, “Improving our alignment and security efforts,” Aug. 31, 2026; OpenAI, “The Hugging Face incident and the road ahead,” Aug. 26, 2026; METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI/Hugging Face hacking incident,” Aug. 26, 2026; UK AI Security Institute, “Cheating behaviour in frontier model evaluations,” July 21, 2026; Daniel Johnson, “From AI to killer robots: UN chief issues urgent governance call,” UN News, July 6, 2026; AI Futures Project, “AI Futures Model: Timelines & takeoff,” accessed Sept. 8, 2026. 

  2. Upen Sachdev, Lynne Challender, Ali Ziaee, Anjali Shaikh, Michael Wilson, and Diana Kearns-Manolatos, “Rethinking the CISO role for an AI-saturated enterprise,” Deloitte Insights, Sept. 3, 2026; Mehdi Houdaigui, Cora Chiu, and Michael Weil, “The new economics of cyber risk,” Deloitte and The Wall Street Journal, May 27, 2026; Dario Amodei, “We must pace the frontier,” September 2026; Anthropic, “Claude Fable 5 and Claude Mythos 5,” June 9, 2026.

  3. National Institute of Standards and Technology, “Artificial intelligence risk management framework: Generative artificial intelligence profile,” July 2024.

  4. Deloitte, “Governing agentic AI through an agent action enforcement layer: AI observability framework for controllable, trustworthy multiagent systems,” accessed Sept. 28, 2026.

  5. Governor of California, “Governor Newsom signs first-in-the-nation AI safeguards to protect Californians, calls on the federal government to do its part,” Sept. 9, 2026. 

  6. OpenAI, “The Hugging Face incident and the road ahead.”

  7. Anthropic, “Improving our alignment and security efforts.”

  8. METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI/Hugging Face hacking incident.”

ACKNOWLEDGMENTS

Editorial (including production and copyediting): Kavita Majumdar, Anu Augustine, Rithu Thomas, and Sayanika Bordoloi

Cover image by: Sonya Vasilieff, Jim Slatton; Adobe Stock

Knowledge services: Rishitha Bichapogu

COPYRIGHT