Skip to main content
Welcome to Deloitte
If we have selected the wrong experience for you, please change it above.

Agentic AI in Banking: Building the Auditable Bank

Why many institutions are approaching agentic AI in the wrong sequence

Executive Summary

  • Agentic AI changes the governance problem: autonomy creates speed, but also new obligations to evidence, explain and control end-to-end decision loops.
  • In our experience, some banks are moving in the wrong sequence: building agents first and trying to retrofit controls later, usually when scaling pressure is at its highest.
  • The right target state is the Agentic Operating Model: a governed, controlled and auditable end-to-end process in which an agent performs a defined role (people, process, controls, data, technology and evidence).
  • To deliver that target state consistently, institutions need an Agentic Development Framework: a repeatable development process that turns operating-model design into technical build requirements and disciplined deployment.
  • Start by designing the Agentic Operating Model, derive your technical build requirements, deploy (including training and change) and provide ongoing assurance that the process, not just the agent, operates correctly.

In many large banks, AML compliance teams review thousands of alerts every working day. Agentic AI could triage that flow, surfacing higher risk cases, filtering false positives and flagging edge cases for human review. The efficiency case is easy to grasp. The harder question, however, is whether the institution could explain, evidence and defend those decisions if challenged.

If the honest answer is “not yet,” that would not make your institution unusual. But it does point to the harder truth that many firms are still some distance from scaling agentic AI safely. This is because they are building in the wrong order.

The central challenge is not simply technological. It also concerns the operating discipline that surrounds autonomy. Put simply, design the Agentic Operating Model first (the governed and controlled target state), then convert that design into build requirements through an Agentic Development Framework. Once in place, deploy with training and change, and finally assure that the process continues to operate correctly, not just the agent.

The next competitive divide may not be between AI adopters and non-adopters. It may rather be between institutions that can orchestrate autonomy safely and those that layer agents onto fragile foundations.

1. Why agentic is different (and why it matters now)

Earlier waves of automation typically followed written rules: “…if threshold X is breached, do Y”. Such systems were more deterministic, more easily auditable and generally easier to explain. Agentic AI is different. These systems can pursue objectives, adapt their behaviour based on context and operate across workflows with greater autonomy. That difference is not semantic. It fundamentally changes the governance challenge.

In the AML context, a rules-based system flags alerts against fixed parameters. An agentic system, in contrast, can draw together signals across fragmented data, assessing the likelihood that a case warrants escalation and routes to prioritise activity with much greater autonomy. The potential outcomes are richer. The oversight requirements are also materially different.

This matters now because the question is no longer whether banks will explore agentic AI. Many already are. The more important question is whether they can scale autonomy without losing trust, and in particular whether each agent can operate as a defined “role” inside an auditable end‑to‑end process, with clear accountability, evidence and controls.

2. What regulators will test: autonomy must be auditable

In the UK, agentic AI does not sit outside the existing regulatory perimeter. It operates within obligations concerning customer outcomes, accountability, resilience and control. In practice, that means autonomy at scale must remain demonstrably auditable: institutions need to evidence how decisions were made, who is accountable, and which controls constrained the outcome.

Regulators require firms to demonstrate fair, comprehensible outcomes and actively avoid foreseeable harms. In an agentic context, that means being able to show, not merely assert, that autonomous decision loops are operating within acceptable, human understandable parameters. Furthermore, accountability cannot disappear into “the model” and named individuals need practical mechanisms to oversee, challenge and constrain agents operating in their name. Elsewhere, critical services must be able to withstand disruptions under severe but plausible scenarios. A system that is opaque in normal operating conditions could therefore be deeply problematic under stress. Model risk discipline matters too: documentation, validation and ongoing monitoring cannot be treated as one-time exercises1.

When a regulator asks who was accountable for an agent-driven decision, pointing to 'the model' will not be enough.

The instinct to treat these requirements as friction is understandable. But it is also a mistake. Institutions that engineer compliance into their agentic architecture from the outset may start slower, but they will accelerate much faster and more safely than those who do not have an Agentic Operating Model. In the longer run, retrofitted controls designed under pressure will, by contrast, usually end up being costlier, of lower quality and leave organisations on the back foot.

3. Why scaling fails: the common failure modes

In many cases, the pattern is familiar. A high value use case – AML triage, credit decisioning, or customer servicing perhaps – is identified. A pilot produces encouraging results. Pressure then builds to scale. And that is when weaknesses surface, often because the Agentic Operating Model was not designed upfront and controls are being retrofitted under time pressure.

In some cases, the data foundation is often not ready for autonomous orchestration in production. Teams may align multiple data sources for a pilot, but those sources are frequently not aligned, timely or governed at scale. Solving data issues one agent at a time increases complexity. A deliberate data strategy and architecture – covering structured and unstructured data, definitions and lineage – pays dividends because it simplifies scaling.

Elsewhere, explainability can too often be treated as a reporting exercise rather than a design principle, one that should be engineered into the system from the outset – shaping what decisions the agent can make, what evidence it must surface and what conditions should trigger human escalation.

And then, the Agentic Operating Model itself can be treated as a post pilot technical task – and too often the focus stays on the agent in isolation. By the time a pilot proves value, stakeholders do not want to wait. If controls, accountability, evidence requirements and escalation paths are not already designed, teams default to retrofitting controls or relying on low level monitoring. The result is a sub standard approach that is not aligned to the organisation’s broader control framework.

4. What to build first: the sequenced approach

Institutions making credible progress separate two things clearly:

  1. The Agentic Operating Model – a governed, controlled and auditable target state
  2. The Agentic Development Framework – the repeatable development process used to design, build, deploy and assure that target state.

Getting sequencing right is essential:

Define the end to end process in which the agent will operate: with accountability and escalation, control objectives, evidence requirements, human machine boundaries, data lineage, resilience and monitoring. Remember, the agent is a role inside this process, not the process itself.

Translate operating model design into a set of technical requirements: permitted actions, prohibited uses, required evidence to surface, audit logs, integration points, guardrails, testing and monitoring. This is where the Agentic Development Framework turns governance intent into buildable specifications.

Treat deployment as process change, not a model release: update procedures, train users, define decision rights, tune escalation thresholds, and ensure front line teams can interpret and challenge outputs.

Provide ongoing assurance that controls and outcomes remain effective in production: monitor performance and drift, test against control objectives, audit evidence trails, and periodically re-validate the end-to-end process. In regulated contexts, assurance must cover the agentified process, not only the agent.

5. The questions leaders should be asking

As institutions look to scale autonomy in a safe and sustainable manner, there are several questions that leaders within these organisations should ask themselves.

  • Where can autonomous decision loops create durable advantage, and where might they create exposures the institution cannot yet govern at deployment speed?
  • Are data quality, strategy, architecture and tooling all sufficient to enable rapid agentic scaling?
  • Is designing the Agentic Operating Model a non negotiable part of the Agentic Development Framework for every use case? Further:
  1. Can accountability for agent driven outcomes be traced clearly to a named owner, with appropriate evidence and escalation paths?
  2. Are control objectives and human understandable evidence requirements intrinsic to the requirements specification?
  3. Are training and process change management part of the standard deployment approach?
  4. Are investments in agentic tools being matched by equivalent investments in controls, testing, monitoring and assurance?

The agentic era is unlikely to reward unbounded experimentation or institutional paralysis. Returns will instead accrue to those who practice disciplined ambition: organisations that design the Agentic Operating Model first, convert it into build requirements through an Agentic Development Framework, deploy with training and change, and then sustain trust through ongoing assurance of the end to end process.

Banking has always been a trust business. In an autonomous era, that trust will depend on how well institutions can design, evidence and sustain control.

References

1. Our UK AI Hub provides space for Deloitte’s topic experts to publish their perspectives on the major transformation questions raised by AI adoption in banking and financial services: how institutions modernise their technology estates, redesign operating models, reshape workforces, strengthen governance and scale AI safely and effectively. Regulation is central to that agenda and is a substantial and rapidly evolving topic in its own right. The regulatory implications of AI depend heavily on use case, institution, type of customer or client affected, the jurisdictions involved and the ways in which technology is designed, deployed and controlled. For that reason, this series of article focuses on the transformation themes at hand. For detailed information on regulatory considerations related to AI, we direct readers to our European Centre for Regulatory Strategy (ECRS) for the latest regulatory developments in AI. https://www.deloitte.com/uk/en/blogs/ecrs.html