In many large banks, AML compliance teams review thousands of alerts every working day. Agentic AI could triage that flow, surfacing higher risk cases, filtering false positives and flagging edge cases for human review. The efficiency case is easy to grasp. The harder question, however, is whether the institution could explain, evidence and defend those decisions if challenged.
If the honest answer is “not yet,” that would not make your institution unusual. But it does point to the harder truth that many firms are still some distance from scaling agentic AI safely. This is because they are building in the wrong order.
The central challenge is not simply technological. It also concerns the operating discipline that surrounds autonomy. Put simply, design the Agentic Operating Model first (the governed and controlled target state), then convert that design into build requirements through an Agentic Development Framework. Once in place, deploy with training and change, and finally assure that the process continues to operate correctly, not just the agent.
The next competitive divide may not be between AI adopters and non-adopters. It may rather be between institutions that can orchestrate autonomy safely and those that layer agents onto fragile foundations.
Earlier waves of automation typically followed written rules: “…if threshold X is breached, do Y”. Such systems were more deterministic, more easily auditable and generally easier to explain. Agentic AI is different. These systems can pursue objectives, adapt their behaviour based on context and operate across workflows with greater autonomy. That difference is not semantic. It fundamentally changes the governance challenge.
In the AML context, a rules-based system flags alerts against fixed parameters. An agentic system, in contrast, can draw together signals across fragmented data, assessing the likelihood that a case warrants escalation and routes to prioritise activity with much greater autonomy. The potential outcomes are richer. The oversight requirements are also materially different.
This matters now because the question is no longer whether banks will explore agentic AI. Many already are. The more important question is whether they can scale autonomy without losing trust, and in particular whether each agent can operate as a defined “role” inside an auditable end‑to‑end process, with clear accountability, evidence and controls.
In the UK, agentic AI does not sit outside the existing regulatory perimeter. It operates within obligations concerning customer outcomes, accountability, resilience and control. In practice, that means autonomy at scale must remain demonstrably auditable: institutions need to evidence how decisions were made, who is accountable, and which controls constrained the outcome.
Regulators require firms to demonstrate fair, comprehensible outcomes and actively avoid foreseeable harms. In an agentic context, that means being able to show, not merely assert, that autonomous decision loops are operating within acceptable, human understandable parameters. Furthermore, accountability cannot disappear into “the model” and named individuals need practical mechanisms to oversee, challenge and constrain agents operating in their name. Elsewhere, critical services must be able to withstand disruptions under severe but plausible scenarios. A system that is opaque in normal operating conditions could therefore be deeply problematic under stress. Model risk discipline matters too: documentation, validation and ongoing monitoring cannot be treated as one-time exercises1.
When a regulator asks who was accountable for an agent-driven decision, pointing to 'the model' will not be enough.
The instinct to treat these requirements as friction is understandable. But it is also a mistake. Institutions that engineer compliance into their agentic architecture from the outset may start slower, but they will accelerate much faster and more safely than those who do not have an Agentic Operating Model. In the longer run, retrofitted controls designed under pressure will, by contrast, usually end up being costlier, of lower quality and leave organisations on the back foot.
In many cases, the pattern is familiar. A high value use case – AML triage, credit decisioning, or customer servicing perhaps – is identified. A pilot produces encouraging results. Pressure then builds to scale. And that is when weaknesses surface, often because the Agentic Operating Model was not designed upfront and controls are being retrofitted under time pressure.
In some cases, the data foundation is often not ready for autonomous orchestration in production. Teams may align multiple data sources for a pilot, but those sources are frequently not aligned, timely or governed at scale. Solving data issues one agent at a time increases complexity. A deliberate data strategy and architecture – covering structured and unstructured data, definitions and lineage – pays dividends because it simplifies scaling.
Elsewhere, explainability can too often be treated as a reporting exercise rather than a design principle, one that should be engineered into the system from the outset – shaping what decisions the agent can make, what evidence it must surface and what conditions should trigger human escalation.
And then, the Agentic Operating Model itself can be treated as a post pilot technical task – and too often the focus stays on the agent in isolation. By the time a pilot proves value, stakeholders do not want to wait. If controls, accountability, evidence requirements and escalation paths are not already designed, teams default to retrofitting controls or relying on low level monitoring. The result is a sub standard approach that is not aligned to the organisation’s broader control framework.
Institutions making credible progress separate two things clearly:
Getting sequencing right is essential:
As institutions look to scale autonomy in a safe and sustainable manner, there are several questions that leaders within these organisations should ask themselves.
The agentic era is unlikely to reward unbounded experimentation or institutional paralysis. Returns will instead accrue to those who practice disciplined ambition: organisations that design the Agentic Operating Model first, convert it into build requirements through an Agentic Development Framework, deploy with training and change, and then sustain trust through ongoing assurance of the end to end process.
Banking has always been a trust business. In an autonomous era, that trust will depend on how well institutions can design, evidence and sustain control.
References
1. Our UK AI Hub provides space for Deloitte’s topic experts to publish their perspectives on the major transformation questions raised by AI adoption in banking and financial services: how institutions modernise their technology estates, redesign operating models, reshape workforces, strengthen governance and scale AI safely and effectively. Regulation is central to that agenda and is a substantial and rapidly evolving topic in its own right. The regulatory implications of AI depend heavily on use case, institution, type of customer or client affected, the jurisdictions involved and the ways in which technology is designed, deployed and controlled. For that reason, this series of article focuses on the transformation themes at hand. For detailed information on regulatory considerations related to AI, we direct readers to our European Centre for Regulatory Strategy (ECRS) for the latest regulatory developments in AI. https://www.deloitte.com/uk/en/blogs/ecrs.html