Featured answer: AI model risk management in banking should govern the complete decision system—not only the foundation model. In 2026, banks need a risk-tiered inventory, use-case validation, data and prompt lineage, tested human oversight, provider-change controls, operational-resilience testing and clear accountability across the EU AI Act, DORA and prudential model-risk expectations.
Generative AI moved quickly from experimentation into customer service, document analysis, coding, financial-crime controls, credit processes and risk reporting. That shift changes the question for banks. The issue is no longer whether a large language model can produce useful text; it is whether an institution can demonstrate that an AI-enabled process remains accurate enough, controlled, resilient and accountable for its intended decision.
The timing is material. The EU AI Act entered into force in August 2024 and, subject to its phased exceptions, became generally applicable on 2 August 2026. DORA has applied since 17 January 2025. Neither framework replaces prudential model-risk management. Banks therefore need a control architecture that reconciles legal classification, ICT resilience and the governance of models used in financial decisions.
Key takeaways
- Classify the complete AI use case, including retrieval, prompts, rules, interfaces and downstream actions.
- Tier control intensity by possible harm, financial materiality, autonomy and reversibility.
- Validate outcomes and failure modes, not only benchmark accuracy.
- Treat human oversight as a control whose effectiveness must be tested.
- Contractual assurance from a provider does not transfer accountability away from the bank.
Traditional model risk and generative-AI risk
Traditional bank models usually transform controlled inputs into comparatively stable numerical outputs. Generative systems may produce different responses to semantically similar prompts, rely on very large training corpora, and change through provider updates that the bank does not control. They can generate plausible but unsupported statements, reveal sensitive data through poorly designed retrieval, or behave differently when users discover unexpected prompt patterns.
The response should extend—not discard—established model-risk principles. An inventory, materiality assessment, independent challenge, change control, monitoring and accountable use remain essential. The difference is the system boundary. A bank does not deploy a foundation model in isolation. It deploys an application containing a model version, system prompts, retrieval sources, filters, orchestration logic, user permissions, human review and interfaces to other systems.
A risk-tiered inventory
Every AI use case should record its purpose, owner, legal classification, affected customers or processes, data sources, provider, model version, deployment pattern, autonomy, downstream actions and fallback. Materiality should consider expected exposure and severity of harm, but also speed and scale. A weak output used by one analyst is different from the same output distributed automatically to thousands of customers.
A practical risk score can combine impact, exposure, autonomy and detectability. For example, a bank might use R = I × E × A × (1-D), where impact and exposure are normalised measures, autonomy captures the extent to which outputs trigger action, and D represents the probability that an error is detected before harm. The formula is a governance device, not a regulatory capital model: its purpose is consistent triage and transparent override.
Validation must match the use case
Validation begins with intended use and prohibited use. Test sets should represent real language, customer groups, product terms and edge cases. The assessment should cover factual grounding, completeness, bias, stability, data leakage, prompt injection, refusal behaviour, explainability appropriate to the user and consequences of false positives and false negatives.
For retrieval-augmented generation, validators should trace answers to approved sources and test what happens when sources conflict, become stale or lack an answer. For agentic functions, testing must include action boundaries, permission escalation, repeated tool calls and recovery from partial failure. For code generation, the control objective is not elegant text but secure, reviewed and reproducible code.
Human oversight is not a label
Calling a process “human in the loop” proves little. Reviewers must receive enough evidence to identify unsupported output, have time to challenge it, and possess authority to stop the process. Monitoring should measure override rates, reviewer disagreement, detected errors and cases in which staff accepted outputs without meaningful review.
Automation bias is a predictable control weakness. A bank should therefore compare performance with and without AI assistance and test whether reviewers become less vigilant when outputs are confident or well written. Where errors are difficult to detect, the model should not support high-impact decisions without stronger independent evidence.
Third-party foundation models and operational resilience
Provider due diligence should cover training and update practices to the extent disclosed, security, data retention, geographic processing, subcontractors, incident notification, service levels, audit rights and exit. The bank also needs technical portability: prompts, evaluation sets and retrieval data should not be inseparable from one provider where the use case supports a critical function.
DORA makes ICT third-party dependence part of the broader operational-resilience framework. AI applications supporting critical or important functions therefore require mapping, continuity arrangements, incident management and realistic recovery tests. A safe fallback may be manual processing, a rules-based system, a secondary provider or suspension of a non-essential function.
Regulatory perspective
The EU AI Act is law and applies through a phased timetable; classification depends on the specific system and use. DORA is directly applicable EU regulation governing ICT risk in covered financial entities. The ECB Guide to internal models describes supervisory understanding of applicable prudential requirements for regulatory internal models; it should not be presented as extending legal scope to every generative-AI tool. Broader AI governance should instead combine applicable law, existing model-risk policy and the institution's own operational and conduct-risk framework.
Hypothetical practical example
Assume a bank introduces an AI assistant that drafts credit-review memoranda from financial statements and internal files. It does not approve credit, but its text may influence the analyst. The bank inventories the full application, restricts retrieval to approved documents, samples outputs against expert-written memoranda and weights missing covenant information more severely than stylistic differences. The control threshold is set so that material factual errors trigger suspension and revalidation. Human reviewers must see source links beside every material statement.
The example numbers and thresholds would be institution-specific. The important design principle is that testing mirrors the decision and assigns loss or harm weights to different failure modes.
What CROs should do now
- Create one inventory across models, AI applications and material end-user tools.
- Define risk tiers using harm, materiality, autonomy, scale and reversibility.
- Require independent pre-deployment evaluation for material use cases.
- Test whether human oversight detects the failures it is intended to control.
- Map AI providers and subcontractors into DORA dependency and exit planning.
- Establish triggers for provider updates, drift, incidents and revalidation.
- Report material limitations and accepted residual risk to accountable executives.
Conclusion
The central governance error is to treat generative AI either as ordinary software or as an unknowable technology beyond established control. It is neither. Banks can govern it by defining the decision system, measuring use-case failures, testing operational dependencies and assigning accountable ownership. The strongest 2026 framework connects AI classification, model risk, operational resilience and customer outcomes without allowing any one label to substitute for evidence.
References
- European Commission, AI Act regulatory framework and implementation timeline
- European Union, Regulation (EU) 2022/2554 on digital operational resilience, 14 December 2022
- ECB Banking Supervision, revised Guide to internal models, 28 July 2025
Frequently Asked Questions
Is generative AI a model under bank model-risk governance?
The answer should depend on use and materiality, not the product label. If an AI component influences a financial, prudential, customer or control decision, it should be inventoried and governed within a defined system boundary.
What makes generative AI validation different?
Validation must assess variable outputs, unsupported statements, prompt and retrieval design, security, bias, provider changes and the effectiveness of human review in addition to conventional conceptual soundness and performance.
How do the EU AI Act and DORA interact?
They address different dimensions. The AI Act is a risk-based legal framework for AI systems, while DORA governs ICT risk and operational resilience in financial entities. A banking AI service may require controls under both.
Can a bank rely on a foundation-model provider's testing?
Provider evidence is useful but not sufficient. The bank remains responsible for validating its own use case, data, prompts, retrieval layer, application controls, users and downstream decisions.