Skip to main content
Welcome to Deloitte
If we have selected the wrong experience for you, please change it above.

AI Workforce Transformation: Why Critical Thinking Matters

The workforce challenge of the agentic era that isn’t being talked about

Executive Summary

  • The primary workforce risk in the agentic era is the erosion of human judgment. As AI outputs become more reliable and frequent, professionals increasingly defer to them, weakening meaningful human-in-the-loop control without any obvious system failure.
  • Human oversight risks becoming “ceremonial” unless actively reinforced. Effective control depends on institutions enabling challenge. If not, declining override rates and questioning will degrade the control environment despite stable model performance.
  • Autonomy concentrates accountability on senior individuals, not systems. To comply with regulation, leaders must demonstrate a real understanding of how systems operate, fail, and are governed. Being present in the loop is not enough.
  • AI is shifting human work toward judgment, escalation and control, raising skill requirements. As routine tasks are automated, roles in risk, compliance and customer operations will require deeper analytical, ethical, and decision-making capabilities. However, it is still not yet clear what the long-term impact of automation will be on junior roles, where future leaders learn their trade and develop the skills of judgement and challenge they will come to rely on later in their careers.
  • Workforce strategy and governance must evolve together, with new capabilities and metrics. Institutions need to AI build skills and track indicators like override rates, challenge quality and escalation effectiveness to ensure oversight is substantive

There is a workforce failure mode in the AI-enabled bank that rarely shows up in model dashboards, operational resilience testing or regulatory returns. Neither does it trigger alerts. It develops quietly, one approval at a time.

Consider, for example, the situation of an experienced analyst who receives an AI-generated recommendation. The model output is confident and the logic appears plausible. The system has also been correct on other recent occasions and, under time pressure, the analyst approves the case faster, with less challenge than they would otherwise have applied before the agentic system was introduced.

This is automation bias in action. In banking, its significance lies in how ordinary it is. No dramatic model failure is required. Judgment simply begins to migrate from the person to the system, until the human-in-the-loop (HITL), while present, no longer exerts meaningful control.

A central workforce risk in the agentic era is the gradual transfer of judgement from trained professionals to systems that appear reliable.

Automation bias: the failure mode hiding in plain sight

Automation bias is well documented in safety-critical industries where automated systems are usually correct and human attention is stretched. For example, in aviation, healthcare and nuclear power generation, where operators have developed detailed protocols over years to counteract the human tendency to defer to automated systems. Banking is now witnessing many of the same conditions, with higher-volume decision environments, compressed timeframes and outputs delivered by machines with the confidence of settled fact.

Under these conditions, challenge can feel unnecessary and inefficient. A reviewer who questions every recommendation will slow a process down. Yet, a reviewer who rarely questions any of them weakens the whole control environment. In this way, human oversight only really matters when institutions expect and enable genuine challenge.

Illustrative failure module

A bank's credit underwriting team uses a supervised agentic system to recommend approval or referral on personal loan applications. Over eighteen months, the team's override rate falls sharply despite the model's accuracy metrics remaining steady. No alert is triggered. A later review finds that analysts had gradually stopped applying independent judgment to edge cases, including a segment where human review had previously caught instances of recurring bias. The issue had been propagating undetected for months.

The compliance implications here are significant. Regulators require accountability for regulated outcomes rest with named individuals, and institutions must actively avoid foreseeable customer harms1. In such cases, institutions must be able to demonstrate that accountable individuals are exercising substantive judgement. Presence alone is not the same as control, and neither obligation can be discharged by a workforce that is simply going through the motions.

If your institution tracks model accuracy but not human override rates, escalation frequency and challenge quality, you are measuring system performance, not the quality of human oversight. Both matter.

Autonomy concentrates accountability

There is a persistent misconception that autonomous systems spread accountability across algorithms, workflows and teams. Yet, in the highly regulated world of banking, the opposite is true. As systems act faster and at greater scale, autonomy concentrates accountability on the individuals who approve systems, define their controls and oversee their use.

That accountability cannot be theoretical. Senior managers need enough understanding of how agentic systems function, how they fail and what meaningful oversight looks like in practice as a means of discharging their legal and governance responsibilities. A senior manager who cannot describe a system’s escalation pathway, oversight boundaries or validation cycle will struggle to evidence effective control when challenged.

Illustrative failure mode

A Chief Risk Officer (CRO) at a mid-tier bank is subject to supervisory scrutiny following regulatory review. The review finds that an agentic customer communications system had generated unsuitable prompts for customers in financial difficulty. The CRO cannot clearly describe the escalation pathway, confirm when the system was last independently challenged nor identify where accountability for oversight formally sits, exposing a material governance gap.

AI changes the shape of work

Across banking functions, there is a broadly consistent pattern. As AI absorbs more repeatable work, human effort moves from execution to judgment, exception handling, escalation and control. The shift is real, but the capability implications differ sharply by function. Specifically:

Risk teams spend less time reviewing outputs in isolation and more time probing edge cases, testing decision boundaries and challenging the logic behind model recommendations. Analysts need to ask why a system reached a specific conclusion, not just whether the conclusion looks plausible, requiring deeper capabilities than those needed for simple output review. 

Compliance teams are moving upstream. Their role will increasingly include shaping control design, testing fairness and evidencing that autonomous workflows remain within conduct and resilience boundaries. In an environment where systems act in seconds, a compliance function that reviews decisions after the fact will arrive too late.

As AI handles more routine interactions, human colleagues are left with the cases where context, vulnerability, discretion and empathy matter most. That raises the skill threshold of the emerging human role and both training and support must reflect this new reality.

A question of displacement (and why avoiding it is a mistake)

An executive team discussing workforce transformation in the context of agentic AI will eventually come to the question of how role volumes, role design and career paths will change. Leaders will gain little by avoiding the matter. Indeed, strong institutions will answer it directly and explain how the transition will be managed.In certain workflows, manual task volumes will fall materially. Some roles will be redesigned, and some teams will become smaller over time. The strategic issues concern how that transition is handled, and how institutions retain the skills and experience needed to supervise agentic systems safely.

Illustrative failure mode

A bank deploys an agentic AML triage system and reduces the number of experienced investigators in its financial crime operations team. Eighteen months later, a complex cross-border case surfaces patterns the model flags with low confidence. The remaining oversight team, now weighted towards junior analysts with limited investigation experience, does not escalate. The case is closed. A later review identifies a missed suspicious activity report (SAR). The institutional knowledge that would once have caught it had already been removed.

This is particularly problematic where entry-level roles are concerned. Many of today’s leaders honed their skills by working their way up through financial organisations, gaining valuable experience along the way. Yet, in the event that entry-level positions – many of which, by their nature, involve more manual and repeatable tasks – become scarcer as the activities within them are automated away, organisations need also concern themselves with where the next generation of middle- and upper-management will come from?

The people best placed to supervise AI-enabled workflows are often those who understand the manual process those workflows were built to replace. Losing them creates operational risk.

Building the skills architecture for governed autonomy

Workforce redesign in the agentic era goes well beyond generic upskilling. Banks need to build stronger institutional capabilities across five key areas: AI literacy, data fluency, ethical reasoning, oversight, discipline and resistance to automation bias. Many institutions are still in the early stages of this journey, and the maturity model below shows the capabilities that exist at each level.

Table 2: AI Governance Capability Maturity Model

Capability

Level 1: Aware

Level 2: Applying

Level 3: Governing

AI Literacy

Understands what agentic AI is; Knows basic failure modes.

Critically assesses model outputs; Identifies anomalies in agent behaviour.

Specifies HITL boundaries; Contributes to governance design.

Data Fluency

Interprets dashboards and standard reports.

Assesses data quality; Recognises bias in outputs; Queries lineage.

Designs data quality standards for agentic inputs; Owns fairness metrics.

Ethical Reasoning

Aware of regulatory requirements; Completes mandatory training.

Applies fairness lens to decisions; Escalates concerns proactively.

Embeds ethical constraints into system design; Coaches others.

Oversight Discipline

Follows HITL protocols when prompted.

Exercises independent judgment; Challenges model outputs routinely.

Designs oversight frameworks; Tests escalation pathways regularly.

Resistance to Automation Bias

Aware of automation bias as a concept.

Actively questions high-confidence outputs; Documents override rationale.

Creates a culture of challenge; Tracks override quality and challenge rates as governance indicators.

Source: Deloitte

The fifth capability, resistance to automation bias, is often the least developed and most directly linked to the risks described above. Awareness training alone will not fix it. Rather, institutions need structural reinforcement, with documented overrides, regular calibration on cases where the model was wrong and management metrics that treat persistently low challenge rates as a governance signal rather than an indicator of efficiency.

Diagnostic question

Review your last twelve months of HITL override data. If override rates are falling steadily and challenge documentation is sparse or formulaic, you may be witnessing automation bias becoming institutionalised. That is both a conduct and governance risk.

The talent you cannot train from within

Training existing staff addresses only part of the capability gap. Banks also need people who can combine a technical understanding of agentic systems with regulatory literacy, control design and operating model judgement. Such talent is scarce and competition for it is intensifying.

Key roles include:

Able to validate agentic decision loops, not just statistical models. These specialists are likely found in a shallow pool spanning quantitative finance, machine learning research and regulatory affairs. The combination is rare, and institutions not building their talent pipeline now will be competing for constrained supply at precisely the moment when they need this capability most.

Talent acquisition for these scarce roles requires banks to look beyond traditional financial services hiring towards technology firms, consultancies with deep AI governance practices and, in some cases, to regulatory bodies where AI oversight experience is being built in real time.

Board and executive education is a separate requirement. Senior leaders do not need to become technologists. They do, however, need enough understanding to challenge how these systems fail, where escalation sits and what effective oversight looks like in practice. That knowledge needs to sit inside the governance infrastructure of any institution deploying agentic capability at scale.

Measuring what the agentic workforce actually does

Performance frameworks built for linear operating models are no longer enough. Throughput and cost still matter, but they do not tell leaders whether human oversight is working. Institutions need measures that would not have appeared on a pre-agentic scorecard. These include:

Finance has an important role to play in embedding these metrics into investment cases. The resilience benefits of a genuinely capable oversight workforce, such as lower remediation exposure, faster regulatory response and stronger supervisory confidence, are material to the value case for agentic deployment. Institutions that can evidence substantive rather than ‘ceremonial’ human oversight will also be able to move faster on future deployments.

The questions leaders should be asking

  • Do we know the override rate, escalation rate and challenge quality for each AI-enabled workflow currently in production?
  • Can the accountable senior manager describe where human judgement sits “in the loop,” what triggers escalation and when the system was last independently challenged?
  • Are workforce transition plans geared to retain those whose process knowledge is most valuable for safe supervision?
  • Do performance measures reward genuine oversight, or only throughput and cost control?
  • Is resistance to automation bias being reinforced through management practices and training, or only treated as a conceptual risk?

The AI-enabled bank will not be judged solely by its technical sophistication. It will also be judged by whether the people overseeing autonomous systems are equipped, expected and permitted to challenge them.

That makes workforce strategy a core governance issue. Institutions need people who can recognise when judgement is slipping from the human to the machine, and step in before oversight becomes ceremonial.

The AI-enabled bank needs people who can govern what themachine does. Building that workforce is a strategic imperative.

References

1. Our UK AI Hub provides space for Deloitte’s topic experts to publish their perspectives on the major transformation questions raised by AI adoption in banking and financial services: how institutions modernise their technology estates, redesign operating models, reshape workforces, strengthen governance and scale AI safely and effectively. Regulation is central to that agenda and is a substantial and rapidly evolving topic in its own right. The regulatory implications of AI depend heavily on use case, institution, type of customer or client affected, the jurisdictions involved and the ways in which technology is designed, deployed and controlled. For that reason, this series of article focuses on the transformation themes at hand. For detailed information on regulatory considerations related to AI, we direct readers to our European Centre for Regulatory Strategy (ECRS) for the latest regulatory developments in AI. https://www.deloitte.com/uk/en/blogs/ecrs.html

Our thinking