Customer service was the first place most banks put generative AI in front of customers, and it is where the gap between the business case and the customer's experience is easiest to see. The spreadsheet case is straightforward: contacts are high-volume, many are repetitive, and every call that a bot contains is a cost avoided. The customer's case is different. Customers want their problem solved quickly, and they want a person when the problem is complicated, stressful or about money they cannot afford to lose.
The evidence from the past two years is now clear enough to draw lessons. Virtual assistants work at very large scale. Programmes designed mainly to cut headcount tend to disappoint. And regulators, led by the UK's FCA, are judging these programmes on outcomes rather than deflection.
Scale is no longer the question
One large US bank reports that 20.6 million users interacted with its AI-powered virtual assistant nearly 700 million times in 2025, and that interactions have passed 3.2 billion since its 2018 launch3. That assistant has been built up over years, starting with narrow, high-confidence tasks and expanding as usage data showed where it was trusted.
A European buy-now-pay-later provider showed what a generative model can do in a single step. In its first month, its AI assistant held 2.3 million conversations, two-thirds of all service chats. The company said this was equivalent to the work of 700 full-time agents, with 25% fewer repeat inquiries and resolution times cut from 11 minutes to under two1. It projected a US$40 million profit improvement for 20241.
What a first month of generative AI service looked like
Reported first-month results for one European payments provider's AI assistant
| Metric | Reported result |
|---|---|
| Conversations handled | 2.3 million |
| Share of service chats | Two-thirds |
| Equivalent workload | 700 full-time agents |
| Repeat inquiries | 25% lower |
| Time to resolve | Under 2 minutes (from 11) |
| Projected 2024 profit improvement | US$40 million |
Note: Company-reported figures, February 2024. The 700-agent figure is a workload equivalent, not reported job cuts.
When containment becomes the goal
The same provider's chief executive later said the company had gone too far. In May 2025 he said that cost had been ‘a too predominant evaluation factor’ and that ‘what you end up having is lower quality’. He added that customers must know ‘there will be always a human if you want’, and the company began reinvesting in human support2.
In August 2025 a large Australian bank reversed a decision to make 45 customer service roles redundant after introducing an AI voice bot. The finance-sector union had argued that call volumes were rising and that the bank was offering overtime and asking team leaders to take calls. The bank said its initial assessment ‘did not adequately consider all relevant business considerations’ and that the roles were not redundant4.
These are not isolated cases. Cost-led plans to take people out of service have repeatedly run into the same problem: customers remain sceptical of AI-only service, and a common worry is that it will be harder to reach a person.
The pattern behind these reversals is easy to miss on a dashboard. A bot that deflects a contact it cannot resolve does not remove the cost. It moves it to a later call, a complaint or a lost customer. Cost per contact falls, but cost per resolved problem can rise. Measured on the wrong unit, a programme can look successful for months while customers and front-line staff absorb the damage.
What the Consumer Duty changes
In the UK, the Consumer Duty has applied to open products since 31 July 2023 and to closed books since 31 July 2024. It sets outcomes for products and services, price and value, consumer understanding and consumer support6. For an AI contact centre, that last outcome is the one that matters. A journey that keeps customers away from resolution, or makes it harder to complain, switch or get help, is a Duty problem however well the bot scores on containment.
The Duty does not prohibit automation. It requires firms to show that automation serves customers. That means testing journeys with real customers, including those with lower digital confidence, monitoring outcomes by segment and acting when the data show harm. A well-designed assistant can help firms meet the Duty by answering instantly at any hour and recognising distress sooner. It can also breach it by trapping customers in loops.
The FCA's Mills Review, published on 6 July 2026, sets the direction. It describes a spectrum of AI autonomy, notes that only one in five UK adults are open to AI making decisions for them, and questions whether a human ‘in the loop’ always provides meaningful challenge. It stresses that ‘better models do not reduce the need for controls; they increase it’5. Firms will need to evidence good outcomes across dynamic, personalised journeys, and that includes customers in vulnerable circumstances5.
When to hand off to a human
The practical design question is the hand-off. We recommend explicit, auditable triggers rather than leaving escalation to the customer's persistence:
- Vulnerability signals, such as bereavement, illness, financial difficulty or distress, route to a trained person immediately.
- Suspected scams or fraud go to a specialist, because speed and judgement both matter and the customer may be under a fraudster's instruction.
- Complaints and disputes are recognised and logged as such, even when the customer does not use the word.
- Repeat contact on the same issue, or a low-confidence answer from the model, triggers escalation rather than another attempt.
- Any request for a human is honoured without a maze.
The banks that get this right treat AI in the contact centre as a supervised workforce. Each agent has a narrow remit, clear escalation rules and a human team accountable for its outcomes. The large US bank's assistant was built that way, one trusted task at a time3. The early reversals were not failures of the technology. They came from trying to take people out of the process before the evidence supported it.