Data Academy · Tutorial 1 of 10 · Banking

Governed AI in banking

Banks have moved from chat assistants for staff to agents that work through credit monitoring, financial crime alerts and customer disputes end to end. Whether those agents scale is decided less by the model than by governance: whether every decision can be traced to a rule, a policy check and a named approver that a supervisor or auditor will accept.

01 What is changing

Where AI in banking is heading

  • From copilots that summarise documents to agents that assemble a case, propose a decision and raise the next step in the bank’s own systems.
  • From scattered pilots to a few domains redesigned end to end, typically financial crime operations, credit monitoring and servicing, with general managers owning the outcome.
  • From one analyst per queue to small teams supervising a set of specialist agents, with human judgement kept for grading, filing and exceptions.
  • From model accuracy as the test to explainability and auditability as the test, as the EU AI Act, DORA and model risk expectations reach agent workflows.

02 Use cases

Three use cases on the governed path

Each use case runs the same path: a business question, governed context, a deterministic rule, specialist agents, a policy check, an action and a record. How the path works →

Illustrative: names and figures are invented to show the flow.

Use case 1

Covenant breach early warning

Commercial borrowers often show stress in their account behaviour and management accounts months before a missed payment. Catching it early means a structured conversation and fewer write-offs instead of a late workout.

“Which of my commercial borrowers are drifting towards a covenant breach, and what should we do about it before quarter end?”Asked by a commercial credit portfolio manager
Context
  • Halden Joinery Ltd holds a £4.2m term facility in the loan system with a net leverage covenant of no more than 3.0x.
  • Q2 management accounts in the credit portal show EBITDA of £1.1m and net debt of £3.6m, a net leverage of about 3.3x.
  • Core banking shows average monthly credit turnover falling from £610k to £455k across the last quarter, down in each of the last two months.
  • The collateral system holds a first charge over the premises, last valued at £2.9m two years ago.
Rule
If reported net leverage is above the covenant level and credit turnover has fallen for two consecutive months, place the borrower on the watch list and open a credit review.
Decision
Move to watch list and review Severity: High
Agents
  • Financial spreading agent Reads the management accounts, recalculates the covenant ratio and flags any adjustments the borrower has made to EBITDA.
  • Cash-flow signal agent Explains the turnover trend from core banking transactions, separating lost customers from seasonal timing.
  • Credit memo agent Drafts the watch-list memo with the breach, the evidence, the collateral position and a stale-valuation note.
Policy
The relationship manager may view the case but cannot change the risk grade. A credit officer must approve the watch-list entry and any grade change, and credit risk must assess whether the change is a significant increase in credit risk for IFRS 9 staging.
Action
A watch-list entry and a credit review task are raised in the loan origination and credit system in a pending approval state, with a request for a fresh collateral valuation.
Data
Loan system: facilities, covenants, repayment scheduleCredit portal: management accounts and compliance certificatesCore banking: account transactions and balancesCollateral system: charges and valuationsRisk grading and IFRS 9 staging history
Use case 2

AML alert triage

Transaction monitoring produces far more alerts than investigators can work properly, and most are closed as explained. Assembling the evidence first lets investigators spend their time on the cases that matter and gives each file a consistent, audit-ready record.

“Which of today’s monitoring alerts need a real investigation, and is the evidence already pulled together?”Asked by a financial crime operations lead
Context
  • Alert TM-88213 in transaction monitoring covers Corvane Trading Ltd: 14 cash deposits of between £9,050 and £9,700 across 6 branches in 10 days, totalling £131,400.
  • The KYC record in the customer due diligence system shows declared turnover of £40,000 a month and a standard risk rating set two years ago.
  • Sanctions and PEP screening returns no matches for the company or its two directors.
  • Case management shows 2 earlier alerts in the last 12 months, both closed as explained by seasonal trade.
Rule
If 5 or more cash deposits between £9,000 and £9,999 are made within 10 days across 3 or more branches, the alert is escalated to an investigator and cannot be closed by an agent.
Decision
Escalate to investigation Severity: High
Agents
  • Evidence assembly agent Pulls the transactions, KYC profile, screening results and prior alerts into a single case file.
  • Network agent Uses the knowledge graph to show linked parties, shared addresses and counterparties of the account.
  • Narrative agent Drafts a factual case summary for the investigator, stating what is known and what is not, without drawing a conclusion.
Policy
Agents may not close, dismiss or file. The decision to make a suspicious activity report rests with the investigator and the nominated officer under the Proceeds of Crime Act 2002, and tipping-off rules mean no customer contact is drafted. Any change to the customer risk rating needs compliance approval.
Action
The case is referred to a human financial crime investigator in the case management system with the evidence pack attached. No account restriction is applied automatically.
Data
Transaction monitoring alerts and scenariosCustomer due diligence and KYC profilesSanctions and PEP screening resultsCase management historyBranch and channel reference data
Use case 3

Card dispute provisional credit

Small card disputes are high volume and customers judge the bank on how quickly they are handled. Clear low-risk cases can be settled within policy straight away, leaving staff for the disputes that need judgement.

“Can we credit this customer today while the chargeback runs, or does it need a person to look at it?”Asked by a card servicing team leader
Context
  • A customer of six years disputes an £84.50 card-not-present transaction with an online merchant, logged in the card dispute system.
  • The card processing system shows the transaction was not authenticated with strong customer authentication.
  • The dispute system shows 1 prior dispute by this customer in the last 24 months, and 37 disputes against the same merchant in the last 30 days.
  • Core banking shows the account in good standing with no arrears.
Rule
If the amount is below £150, the customer has no more than 2 disputes in the last 12 months and the transaction is card-not-present, post a provisional credit and raise a chargeback.
Decision
Provisional credit and chargeback Severity: Low
Agents
  • Dispute intake agent Classifies the dispute reason from the customer’s message and matches it to the card scheme reason code.
  • Merchant pattern agent Notes the cluster of disputes against the merchant and flags it to the merchant risk team.
  • Customer reply agent Drafts a plain-language confirmation of the credit and what happens next.
Policy
Within the servicing team’s automatic authority under the dispute policy. Where the customer reports the payment as unauthorised, the refund timing in the Payment Services Regulations 2017, which implement PSD2, applies. Amounts above the limit or repeat disputers go to a dispute specialist.
Action
The provisional credit is posted in core banking and the chargeback raised in the card dispute system, released automatically within policy and recorded with the rule that applied.
Data
Card dispute system: disputes and reason codesCard processing: authorisations and authentication dataCore banking: account status and postingsCustomer profile and dispute historyMerchant risk data

03 The foundation

What the agents need to understand

Core entities in the ontology

CustomerAccountFacilityCovenantCollateralTransactionCounterpartyAlertCaseRisk Rating

Systems they come from

Core banking
accounts, balances, postings and standing orders
Loan origination and servicing
applications, facilities, covenants and repayment schedules
Card processing and disputes
authorisations, chargebacks and reason codes
Transaction monitoring and case management
alerts, scenarios, investigations and outcomes
Customer due diligence
KYC profiles, beneficial owners and screening results
Collateral management
charges, valuations and security cover

04 Guardrails

The controls that let it scale

1

Rules decide, agents explain

Grading, escalation and refund decisions come from written rules a model risk team can validate, not from a language model’s judgement.

2

Approval for sensitive actions

Risk grade changes, account restrictions and anything touching a suspicious activity report are created as pending approval for a named officer.

3

Explainable credit decisions

Every credit-related outcome carries the rule, the inputs and the data sources used, so it can be explained to the customer and to the supervisor.

4

Data minimisation

Agents see only the fields a case needs, in line with GDPR, and personal data stays within the bank’s controlled environment.

5

Operational resilience

Agent workflows and their third-party dependencies are recorded and tested as ICT services under DORA, with a manual fallback for each.

05 Rollout

From the first use case to many

  1. 1

    Pick one domain with a clear rulebook

    Start where written policy already exists, such as financial crime alert triage or card disputes, so the rules can be encoded and checked.

  2. 2

    Model the business before the agents

    Build the ontology of customers, accounts, facilities and cases and connect the systems that hold them, so agents work on governed entities rather than raw tables.

  3. 3

    Run alongside the existing team

    Let agents prepare cases while people still decide, and compare outcomes until the team trusts the rules and the evidence.

  4. 4

    Open automatic release within limits

    Allow automatic action only for low-risk outcomes within written authority, keeping everything else pending approval.

  5. 5

    Extend to the next domain on the same model

    Reuse the same entities, policies and audit trail for credit monitoring and servicing rather than starting again.

06 What to measure

Outcomes, not activity

Alert handling timeInvestigator time on genuine casesEarly-warning lead time before defaultDispute resolution timeDecisions overturned on reviewAudit findings on AI-assisted decisions

07 Pitfalls

What usually goes wrong

  • Automating the old process. Adding an agent to each step of an unchanged workflow keeps the hand-offs and the queue; redesign the flow around what the agent prepares and what the person decides.
  • Letting the model make the call. When a language model decides who is escalated, the outcome cannot be validated; keep the decision in an explicit rule and use the model to explain it.
  • Weak data foundations. Agents built on inconsistent customer and account records repeat the inconsistencies at speed; reconcile the core entities first.
  • Pilots without an owner. Pilots run by technology teams alone rarely reach production; give a business owner the target and the authority to change the process.

08 Diagnostics

Questions to ask your team

  1. 1

    Which decisions in this domain are governed by a written rule today, and which depend on individual judgement?

  2. 2

    Can we show a supervisor, for any agent-assisted decision, the rule, the data and the person who approved it?

  3. 3

    Which actions are agents allowed to take without approval, and who signed that list off?

  4. 4

    Do our customer, account and facility records agree across core banking, lending and financial crime systems?

09 Keep going

Related reading

— Questions

Frequently asked

Will agents replace credit officers and investigators?

No. In this approach agents assemble evidence and propose an outcome under a rule, while grading, filing and exceptions stay with named people. The work shifts from collecting information to judging it.

How does this fit with model risk management?

The deterministic rules can be validated like any other policy, and the agents’ role is limited to reading, assembling and drafting. Each run is recorded, so model risk and internal audit can review what was used and what was decided.

Where should a bank start?

With one domain that has a clear rulebook and a measurable backlog, such as alert triage or card disputes. Building the business model of customers, accounts and cases there makes the next domain faster.