The next phase of AI will not be defined by faster drafting or better summarisation. It will be defined by whether organisations can trust autonomous AI agents to operate safely across critical business processes.
Over the past two years, many institutions have focused on deploying GenAI through Microsoft Copilot and standalone large language model (LLM) applications. These tools can improve individual productivity, but they often remain “side-of-the-desk” assistants: useful for helping people complete existing tasks more efficiently, but less likely to reshape how work gets done or materially change the return on investment equation.
Firms are increasingly moving beyond point solutions towards an agentic operating model, in which multiple AI agents collaborate autonomously, interact with enterprise systems, make decisions, invoke tools and execute end-to-end business processes with limited human intervention. This is where the potential for more meaningful ROI emerges: moving beyond drafting or summarisation, towards the reconfiguration, automation and optimisation of entire workflows. However, by seizing this opportunity organisations must also fundamentally change how they think about governance, risk and assurance.
Traditional AI assurance approaches assume a relatively static environment, where system behaviour can be documented, reviewed and tested periodically. Agentic systems challenge that assumption, requiring a multidisciplinary response.
When autonomous agents make thousands of decisions each day, assurance can no longer depend on periodic snapshots of system behaviour. In this environment static documentation quickly becomes outdated, point-in-time testing provides only limited confidence and interview-based evidence loses much of its evidential weight.
Assurance must therefore become continuous.
Whilst the principles of trustworthy AI governance remain the same, how those principles are implemented must evolve at pace. Without effective governance, cyber security, privacy, resilience, accountability, human oversight and regulatory compliance, organisations will struggle to scale interconnected agentic ecosystems safely.
For institutions, telemetry and observability are becoming central to their response. Indeed, for many leading organisations, they are increasingly becoming the primary source of evidence for continuous assurance. As a result, many leading firms are investing heavily in observability platforms and telemetry frameworks.
Rather than relying on periodic assessments or static documentation, observability platforms continuously collect evidence concerning how AI systems behave in production. Prompts, tool invocations, model interactions, policy decisions, escalations and human interventions are all becoming part of the assurance evidence base. This gives organisations a richer view of not only what an AI system did, but why it did it and whether it remained within approved governance boundaries.
This represents an important evolution in AI governance. Production monitoring is not new. MLOps has long enabled organisations to track model performance, drift and reliability post-deployment. Agentic AI expands the assurance lens further, requiring organisations to also have visibility of agent goals, reasoning traces, tool use, system interactions, decision pathways, hand-offs between agents and the outcomes of actions taken in live business processes.
Hence, observability is no longer just an engineering capability. In an agentic ecosystem, it becomes a core governance and assurance mechanism also, helping organisations understand how AI systems behave in production, demonstrating that they remain within approved boundaries and allowing institutions to intervene rapidly when behaviours or outcomes diverge from expectations.
Perhaps the most significant shift in assurance is not in how assurance is performed, but in where assurance is intended to provide confidence.
Historically, assurance activities focused on validating the behaviour of individual AI systems and models, seeking to answer questions such as: “Does this model operate correctly, reliably and within acceptable risk tolerances?” The primary challenge for assurance teams was developing testing approaches capable of accurately representing system behaviour and identifying potential failures.
As organisations scale the adoption of AI and agentic systems, that focus is no longer sufficient. The assurance objective expands beyond individual models to the governance structures that oversee them. The key question then becomes: “Can this organisation demonstrate that it governs AI effectively, consistently and safely at enterprise scale?”
The role of assurance is therefore shifting from providing confidence in individual AI components to assessing whether the organisation has the processes, controls, technologies and oversight needed to manage a complex ecosystem of models, agents, tools, workflows and human decision-makers.
Future assurance will therefore focus on the ability of institutions to maintain a comprehensive AI inventory and classify AI use cases according to risk. To manage this risk, they will also require effective governance workflows, an appropriate level of embedded appropriate human oversight, and the deployment of monitoring and observability capabilities to provide ongoing visibility across risk and performance. In this way, the question moves beyond whether a singular system behaved as expected at specific a point in time, but to whether the organisation can continuously govern, monitor and control its AI estate.
Organisations are already using AI to automate aspects of testing, validation and control monitoring. Automated continuous testing, automated policy validation, AI-driven monitoring and real-time control assessments are becoming increasingly common across large enterprises. This automation will be essential as organisations move towards tens of thousands of agentic AI use cases, where traditional testing and validation methods alone would demand unsustainable levels of resource.
As these capabilities mature, and the emergence of fully agentic organisations becomes reality, the role of the first line of defence will also evolve.
Business and technology teams will become increasingly responsible for designing, implementing and operating their own AI governance frameworks, supported by automation and observability platforms that monitor AI systems continuously throughout their lifecycles.
This will also change the role of second-line, third-line and external assurance teams. Rather than re-performing technical testing already executed through automated control frameworks, assurance functions will increasingly provide independent confidence over the effectiveness of the governance framework itself.
This represents a subtle but profound evolution in the role of trustworthy AI assurance.
Of course, automation does not eliminate the need for human judgement.
These changes, and the introduction of agentic systems designed to reduce human friction, require a different approach to when and where human judgement should intervene, and what it means to keep humans in the loop, on the loop or over the loop.
This requires a mature understanding of the guardrails applied to agentic systems especially concerning what agents can access, how they can act and which activities still require manual controls. AI governance itself should never become fully autonomoussince even automated controls still need clear human accountability for configuration, performance thresholds and guardrail definitions.
Some aspects of AI governance are unlikely ever to become fully autonomous, and others should not. Risk acceptance for high-risk AI use cases, ethical judgement, regulatory interpretation, accountability for business decisions, ownership of AI systems and oversight of manual controls all remain fundamentally human responsibilities.
Likewise, independent challenge remains a critical safeguard.
No matter how sophisticated automated governance becomes, organisations will continue to require independent assurance over the human decisions and governance processes that sit outside the automated control environment. The more complex the governance framework, the greater the need for specialist AI assurance to confirm that it operates as intended while maintaining data privacy, security and resilience.
As organisations continue to mature, we believe the assurance market itself will begin to diverge.
Leading organisations with sophisticated engineering capabilities will increasingly automate their testing and monitoring with their primary requirement becoming independent assurance that their own AI governance and automated testing operating models have been appropriately designed and implemented and are operating effectively.
That independent assurance will extend across governance processes, engineering implementation, telemetry quality, data integrity, observability platforms, cyber security, human oversight and the ongoing effectiveness of automated controls.
At the same time, however, many organisations, particularly those earlier in their AI journey or deploying highly complex or high-risk use cases, will continue to require specialist support performing targeted technical testing and validation activities.
Hence, rather than reducing demand for assurance, automation merely changes where assurance creates value, increasingly moving away from executing tests and towards providing trusted confidence that the governance ecosystem itself can be relied upon.
Finally, the regulatory landscape will continue to reinforce this direction of travel.
Whilst today's ecosystem is supported by emerging frameworks such as the NIST AI Risk Management Framework, ISO/IEC 42001 and the EU AI Act, there remains no universally adopted assurance methodology for GenAI or Agentic AI.
As these technologies mature, we expect assurance standards, certification schemes and common control frameworks to become significantly more developed. This will further accelerate the transition towards continuous, standards-based AI assurance supported by telemetry and observability.
The direction of travel is clear. The race to implement agentic AI is accelerating, but governance capability, assurance maturity and organisational trust are struggling to keep pace. That gap exposes adopters to serious operational, regulatory and reputational risk events.
The organisations that create the greatest long-term value will not necessarily be those that deploy AI fastest. They will be those that can demonstrate to boards, regulators, customers and shareholders that their AI systems operate safely, responsibly and within clearly defined governance boundaries.
In that world, assurance becomes far more than a compliance activity. It becomes a strategic capability that enables organisations to scale AI with confidence.