A financial services firm does not need to wait for an AI-specific rulebook before it can investigate a use. It does need to identify which existing requirements apply, who is accountable and what evidence would support safe operation.
The FCA’s current approach to AI says it does not plan extra AI regulations and will rely on existing frameworks, including Consumer Duty and accountability and governance rules. The page describes a principles-based, outcomes-focused approach. That is not approval of a use case. Firms differ in permissions, customers, services and applicable rules.
This article provides three candidate investigations. A useful quarterly result may be a controlled test, a governance decision, prerequisite data work or a conclusion that the use is unsuitable. Legal, compliance, risk, privacy, security, model and professional specialists should interpret current requirements for the firm.
Companion articles apply the same discipline to law firms and management consultancies.
1. Evaluate monitoring triage against an established control
AI or machine-learning systems can help prioritise transactions, communications or activity for review. The opportunity is better use of specialist attention. The risk is that the system misses material behaviour, generates excessive false alerts, embeds bias or obscures why an item was selected.
Automated monitoring is not inherently the lowest-risk AI use. Its consequence depends on the control, decisions, data and people affected.
Select one defined task
Choose a control with an understood purpose, current procedure and measurable baseline. Examples might include prioritising alerts already generated by a rules-based process or classifying a known category for human review. Avoid beginning with a promise to “monitor everything”.
Document the current sampling or triage process, known limitations, error consequence and escalation. Involve the control owner and people who perform the work.
Build a representative evaluation
Use approved historical or safely prepared data. Preserve important rare cases and changes over time. Establish reference decisions with qualified reviewers and record disagreement.
Evaluate missed material cases, false alerts, performance across relevant groups or activities, explanation, reviewer time and how behaviour changes after introduction. A falling false-positive rate is not automatically progress if detection also weakens.
Keep control ownership human and explicit
Define who can change thresholds or models, who investigates alerts, what evidence is retained and who decides the system is no longer fit. Monitor drift and supplier change. Design fallback if the service is unavailable.
The test should examine the combined process, including human workload and escalation, rather than treating model accuracy in isolation as the control outcome.
2. Assess relationship signals without turning communication into surveillance
Changes in meeting attendance, service use or response patterns can prompt a relationship manager to ask whether a client needs support. Automated sentiment analysis of email and other communication creates substantial uncertainty and privacy risk. Tone is contextual, model inference can be wrong and an employee or client may reasonably object to being profiled.
Begin with the service problem, not an assumption that a model can detect “relationship health”.
Use less-intrusive evidence first
Consider information already collected for operating the service: missed agreed reviews, unresolved requests, repeated access failure or an explicit change in communication preference. Aggregated journey evidence may reveal friction without profiling an individual.
If individual-level prompting remains justified, define the specific signal, intended support, person who sees it and action they may take. Present the signal as uncertain context, never as a conclusion that a client will leave or is dissatisfied.
Establish the legal and ethical position
Assess lawful basis, purpose compatibility, transparency, necessity, data minimisation, retention, security, client confidentiality and rights with qualified advisers. A data protection impact assessment may be required depending on the processing and risk; specialists should determine that rather than relying on a generic article.
Consider employee data too. Analysing communications can affect staff as well as clients and may require consultation and employment advice.
The ICO’s AI and data protection guidance provides current regulator material. Parts may be updated as law and policy change, so check the live guidance.
Evaluate helpfulness and harm
Test offline before any live prompt. Include different communication styles, languages, sparse histories and situations where reduced contact is expected. Measure false concern, missed known issues, relationship-manager interpretation, client response, complaints and any unequal effect.
The firm should be prepared to decide that free-text sentiment analysis is too intrusive or unreliable and use a simpler service trigger instead.
3. Generate controlled first drafts of routine reports
Periodic reporting can combine structured data, approved explanatory text and professional commentary. AI may help assemble or draft parts of that output. It cannot take responsibility for accuracy, suitability, fair presentation or the obligations attached to the communication.
Start with one report whose sources, template and review are relatively stable. Keep bespoke advice and high-judgement commentary outside the first test.
Standardise the input before generating text
Identify the authoritative data source, reconciliation, calculation owner and cut-off. A language model should not repair inconsistent numbers or choose between conflicting systems.
Separate deterministic content from generated content. Figures, disclosures and fixed statements may be safer through controlled templates. Use AI only where variation adds value and can be checked.
Design review and provenance
The reviewer should see source data, approved text, generated draft and changes. Define the checks required before release and retain an appropriate audit trail. Prevent automatic dispatch.
Test factual accuracy, unsupported statements, omission, tone, readability, accessibility and consistency across similar clients. Include adverse market conditions and missing data, not only routine periods.
Measure the complete production process
Compare drafting, reconciliation, review, correction and approval time with the baseline. Track error and near miss, reviewer reliance, client questions and control exceptions.
If review consumes most of the process or generated variation creates extra uncertainty, template and workflow improvement may be the better investment.
Build common governance once, adapt it per use
The FCA’s June 2026 explanation of its approach says existing frameworks, including Consumer Duty, SM&CR and governance and controls, remain relevant. The Consumer Duty applies in its defined scope and should not be invoked as a universal rule for every firm or customer.
Create a use-case record covering:
- intended outcome and people affected;
- accountable business and senior owners;
- model, supplier and version;
- data, purpose, permission and flow;
- current process and baseline;
- validation and limitations;
- human decision and escalation;
- consumer or market harm scenarios;
- security, resilience and third-party dependence;
- monitoring, incidents and change control; and
- exit or fallback.
Governance should scale with consequence. A drafting assistant using public, approved text and a system prioritising financial-crime alerts should not pass through the same lightweight review.
Where Consumer Duty applies, use its outcomes as part of product, communication, support and value assessment. Do not claim AI supports the Duty merely because it is faster. Evidence the outcome for relevant customers, including those with characteristics of vulnerability where applicable.
Our related article, AI in financial services: what has cleared compliance and what has not, should be read with current official sources because regulatory and supervisory material evolves.
Choose by readiness and possible harm
Do not rank these three using a universal risk ladder. Score the actual proposal against:
- material problem and current baseline;
- applicable obligations;
- data fitness and permission;
- harm if output is wrong or missed;
- ability to validate rare and important cases;
- human capacity and authority;
- supplier and integration risk;
- reversibility and fallback; and
- quality of the evidence available within the quarter.
Select one bounded evaluation. Agree continue, revise and stop conditions before testing. Report inconclusive results and preserve the decision trail.
If you want a structured assessment of which of these three use cases is the best fit for your firm's current compliance posture, data environment, and operational capability - and what a properly scoped pilot design looks like - book an AI opportunity assessment for financial services firms. We've also built a financial services AI use case scorecard - a one-page readiness assessment covering data availability, regulatory posture, technical environment, and governance maturity for each of the three use cases. It takes about ten minutes and it'll tell you where to start.
The firm that makes one supported decision this quarter has progressed. The comparison that matters is with its own previous evidence and control, not an unsupported claim that competitors are already twelve months ahead.



