Insights from the 2025 EMEA Model Risk Management Survey: How banks and insurers are validating AI models, and the challenges they face.
The adoption of AI continues to accelerate across the financial services sector. As banks and insurers increasingly deploy AI components and solutions, they face a complex landscape of emerging risks spanning operational, legal, conduct, data, cyber, and model risk domains. Insights from Deloitte’s 2025 Model Risk Management survey reveal that although institutions recognise AI as strategically important, their approaches to identifying and validating the associated risks are still maturing.
This article explores the implications from a Model Risk Management (MRM) perspective, focusing specifically on how institutions can identify and mitigate model risks arising from AI use cases, and how to address the evolving challenges this presents. Moreover, it is suggested that while MRM's responsibility is validating the AI component, effective validation requires collaboration with other teams to address non-technical risks arising in AI solution architecture, such as user literacy, workforce adoption, change management, accountability, and automation bias, which might fall outside MRM's scope, but are essential to overall AI governance.
The survey results indicate a clear direction of travel: AI and ML techniques are now widely used across banks and insurers, and GenAI tools, such as chatbots and coding aids have moved from experimentation into broader business use. However, validation activity has not developed at the same pace pointing to an emerging gap between AI adoption and control of the model risk arising as a consequence.
Model risk, at its core, is straightforward: it is the risk of making a wrong decision based on a model's output. This fundamental risk applies equally to traditional statistical models and to AI systems. Our survey findings show that institutions are more likely to validate AI-enabled use cases when these can be absorbed into existing model risk management processes and traditional validation cycles. However, AI introduces new dimensions of complexity that make this risk harder to identify and manage. The challenge becomes clearer when distinguishing between AI components (the underlying algorithms where model risk resides) and broader AI solutions, which encompass complete systems with architecture, integrations, and supporting infrastructure. Traditional MRM framework designs could absorb the validation of AI components. By contrast, GenAI and application-based AI, for example chatbots, customer and marketing tools, or Agentic AI components, require enhanced validation pathways. These AI solutions are often developed or deployed directly by business units and may combine several components, such as GenAI interfaces, retrieval layers, database connectors, and third-party services. As a result, they may not be captured by conventional model inventories or validation triggers.
The implication is that validation practices need to follow the risk, not the label. Where AI use creates model risk, institutions should be able to demonstrate that the risk has been identified, assessed and validated through a proportionate process. For GenAI and more complex AI applications, this may require validation to extend beyond the underlying algorithm and consider the end-to-end system, including data flows, retrieval accuracy, user interaction, controls, monitoring, and governance. Without this broader view, institutions risk leaving AI-related model risk insufficiently assessed, especially as AI adoption continues to accelerate.
Banks and insurers face three main challenges affecting AI validation: technical complexity, outdated validation frameworks, and a complex evolving regulatory landscape.
The first challenge is technical: as AI systems are fundamentally more complex and harder to interpret than traditional models, understanding how these systems arrive at their decisions is increasingly difficult. This opacity makes it harder to assess whether the AI models are working as intended, and more importantly ensuring that AI systems do not discriminate or produce skewed outcomes. The nature of these technical challenges varies by AI type. Traditional machine learning models can fit relatively well into existing validation processes, though they still require careful attention to bias and performance testing. Continuously learning AI systems, however, change faster than traditional validation cycles can capture. For example, AI agents might use feedback loops to learn from the environment they themselves influence, making it difficult to evaluate performance.
The second main challenge is methodological, as traditional validation frameworks are insufficient to cover the new sources of risk arising from AI solutions and to determine appropriate technical methods for validation. Additionally, GenAI systems often exhibit non-traditional lifecycles, moving through iterative development cycles (Proof of concept and minimum viable product stages), so traditional points of validation do not map directly to the points in the AI lifecycle at which risk needs to be validated.
Lastly, the regulatory landscape for AI validation has become increasingly complex, especially as guidance is still being formulated. Banks and insurers are struggling to reconcile how to apply existing regulatory requirements (e.g., SR 11-7/SR 26-2, SS1/23, ECB guidance or Solvency II) to AI, while simultaneously adhering to AI-specific requirements set out by AI regulation such as the EU AI Act, which does not address the banking and insurance specific context. This challenge is compounded by the convergence of multiple overlapping EU regulations in 2026, including DORA (Digital Operational Resilience Act); AMLR (Anti-Money Laundering Regulation); and GDPR (General Data Protection Regulation) — each imposing distinct governance, risk management, and data protection obligations that financial institutions must integrate into a compliance framework for AI.
Different AI types require tailored validation approaches. Full AI validation requires a shift that extends beyond model performance to encompass the entire system architecture. This includes validating user interfaces for transparency, API[1] integrations with other systems, and database governance. For example, for a GenAI chatbot, ensuring the model’s answer quality alone is insufficient, we also need to ensure accuracy in retrieval of relevant information from the connected database, and accuracy recognising the intent behind the user request. Similarly, Agentic AI requires a broader view as system integration, read-and-write permissions, feedback loops and security architecture need to be validated.
Further, a shift from typical validation cycles to increased monitoring is required to adjust to continuously learning AI systems. This approach enables timely detection of performance issues, data drift and emerging risks for proactive management, where models or model context may shift faster than the planned validation intervals.
Finally, integrating vendor-supplied components necessitates a prudent validation approach, due to limited transparency of third-party systems. Vendor agreements should include Service Level Agreements (SLAs) mandating validation evidence to compensate for limited direct access to third-party systems. Moreover, during procurement, the purchasing organisation should conduct light validation assessments (e.g., sandbox testing) to consider as part of the purchasing decision. Identified information gaps from the vendor validation, which represent residual risks, should then be included in post-deployment monitoring.
Despite the new risks generated by AI adoption, AI likely remains one of the most critical sources of competitive advantage in the future. To effectively perform MRM, proper investment in structural expertise and additions to the validation frameworks need to take place. Limited availability of AI validation expertise is a significant factor contributing to current challenges in the validation process. Investing in specialised AI expertise, for example a separate team, is recommended, as AI is no longer just a different modelling technique, but a whole domain with frequent technical and non-technical changes and high levels of complexity in and of itself.
Procedures within the institution should also be appropriate to accommodate GenAI and Agentic AI. Establishing a combined inventory that includes all AI applications is necessary to ensure no models fall outside the scope of oversight. Developing a flexible, AI-specific validation framework, conceived as a living document, will allow validators to continuously update risk assessments and methodologies as the AI landscape evolves. Finally, banks should, as always, look beyond regulatory requirements by exploring industry and academic best practices and fostering internal communities of practice to share knowledge and refine validation approaches. These steps will help build a resilient validation environment that supports risk identification and effective management.
We observe that for many clients creating a dedicated AI Centre of Excellence,[2] a centralised hub that allows cross-leveraging of resources, expertise, and technologies, is an effective solution that also allows for better collaboration between MRM expertise and neighbouring risk and compliance teams involved in AI. Building AI validation capability cannot be MRM's responsibility alone, since the introduced risk spreads across many interconnected areas. Effective validation requires collaboration of cross-functional teams. Specifically, IT and cybersecurity teams must validate system architecture and security; data governance teams must ensure data quality and retrieval accuracy; change management and operational risk teams must address user literacy and workforce adoption; legal and compliance teams must navigate the regulatory landscape and should be aligned on the observed issues and findings.
For the full insights around the validation of AI in banks and insurers, read the survey report here.
[1] Application Programming Interface (API) refers to software interfaces which allows to connect several pieces of software to each other, e.g. the AI component to the main application.
[2] Deloitte, 2026. Leveling up the AI center of excellence.