Are you in control of your AI? Audit it | GRT Consulting

Are you in control of your AI? Audit it | GRT Consulting

Most firms can point to an AI pilot. Fewer can point to an AI audit file that would survive a senior manager interview.

That gap matters. The FCA’s published position is clear: it does not plan a separate AI rulebook. Existing frameworks — Consumer Duty, SM&CR, SYSC systems and controls, outsourcing and operational resilience — already apply to how you design, buy, run and oversee AI. See AI and the FCA: our approach.

The Mills Review (July 2026) makes the same point for retail finance. As autonomy rises, so does the need for named accountability, continuous assurance, and evidence you can reconstruct after the fact.

So the board question is not “do we have an AI policy?” It is: are you in control?

Audit AI regularly — not once at go-live

AI changes after you switch it on. Models drift. Vendors push updates. Staff invent new prompts. Shadow tools appear on phones. A one-off “AI risk assessment” at procurement is not control; it is a snapshot.

Regular audit means a planned cycle that matches how risky and how autonomous the use case is:

Cadence Typical use What you re-check
Continuous / weekly sampling Customer-facing agents, credit, advice, payments, financial crime decisions Outcomes, overrides, incidents, drift alerts, prompt/tool changes
Quarterly Material internal decision support tied to an important business service Performance vs approval pack, data lineage, access rights, third-party changes
Annual (or on material change) Lower-risk drafting aids with strong human gates Inventory accuracy, owner attestations, policy fit, retirement candidates
Triggered Any incident, vendor outage, model swap, new data class, new write access Full re-approval path before scale resumes

If you cannot say when each live use case was last independently reviewed, you are hoping rather than governing.

What areas should you look at?

Keep the review practical. You are not writing a thesis on neural nets. You are testing whether the firm can defend the outcomes.

# Area Why it sits in the audit file
1 Inventory & shadow AI You cannot control what you have not listed — including unsanctioned ChatGPT-style tools
2 Purpose, autonomy & human gates What the system may do alone vs what a person must approve
3 Data classes & leakage paths Customer, market-sensitive, IP and staff data in prompts, logs and vendor stores
4 Model / vendor change control Who may update weights, prompts, tools, plugins; how you detect silent changes
5 Performance, bias & Consumer Duty outcomes Does it still do what you approved, fairly, for the customers you serve?
6 Explainability & evidence Can you reconstruct a decision for a complaint, SAR or SMF interview?
7 Access, identity & over-privilege Agents with write access are operational risk, not “IT hygiene”
8 Resilience & concentration Same model behind several important business services
9 Financial crime & conduct misuse Prompt injection, social engineering, policy invention, unsuitable recommendations
10 Exit & kill switches Can you disable a use case or vendor path without taking the firm offline?

For banks and designated investment firms, the PRA’s SS1/23 model risk management principles remain the cleanest supervisory language for inventory, governance, independent validation and mitigants. That still holds when you stretch the spirit of those principles to generative tools that were not the original use case.

FCA solo firms should not pretend SS1/23 binds them if it does not. They should still recognise the direction of travel.

The PRA’s November 2025 AI/ML model-risk roundtables underline the same themes for dynamic models: more frequent monitoring, quantitative triggers, challenger and fallback options, and kill switches.

Use the three lines of defence you already have

You do not need a fourth line called “AI”. Credit, outsourcing and operational resilience already run through first, second and third line. Put AI through the same machinery.

Line Job on AI
First line — business, ops, product Own the use case: keep it on the inventory, name an owner, write the operating procedures, sample outputs, raise incidents and drill the kill switch
Second line — risk, compliance, financial crime Challenge and set the standards: risk tier, policy, Consumer Duty and SM&CR overlays, monitoring thresholds and open-issue tracking
Third line — internal audit Give independent assurance: risk-based AI audits across the live inventory, not a one-off “AI project review” and then silence

AI can help all three lines. It does not replace them. The Mills Review is blunt on that point. Line 1 still owns the outcomes. Line 2 still challenges. Line 3 still reports independent assurance to the board and audit committee.

What usually goes wrong: Line 1 “owns” a vendor chatbot, Line 2 published principles in 2024, and Line 3 has never put AI into the audit universe. That is not three lines of defence. That is a gap with stationery.

What external support do you actually need?

Not every use case needs a Big Four model validation army. Materiality does.

Bring in independent / external assurance when:

Trigger
The model or agent can commit the firm — customer communications, pricing, credit, advice, payments or regulatory reporting
You rely on a third-party foundation model or managed “AI desk” and need audit rights, evaluation evidence and concentration analysis you cannot generate alone
Internal validation lacks depth on generative systems — prompt injection, tool-use abuse, evaluation harnesses, red-teaming
The board or audit committee wants an outside view before scale — or after an incident

Independent work typically covers:

Scope
Inventory completeness and design effectiveness of controls
Validation of material models and sample testing of outcomes
Vendor due diligence, contractual audit rights, and whether SM&CR Statements of Responsibilities match who can change the system

GRT’s view is simple: buy independence where impact is high; build routine assurance where volume is high. Do not pay for theatre on a meeting-summariser, and do not leave a customer-facing agent with a policy PDF as its only control.

The cost of not auditing

Weak AI oversight is not a theoretical “ethics” problem. It shows up as cash, customers and careers.

Failure mode What it looks like in practice Why audit would have helped
Data leakage Staff paste code, customer files or meeting notes into consumer AI tools — as reported in the Samsung ChatGPT incidents that led to tighter bans Inventory, approved tools, DLP, training evidence, sampling of prompts
Invented policy / improper usage A customer-facing bot invents rules the firm never approved — the Air Canada chatbot case made the firm own the output Scope limits, human gates on commitments, output monitoring vs source of truth
Model-led commercial failure Algorithms that misprice risk at scale — Zillow’s iBuying unwind is the textbook non-FS warning on trusting unchallenged forecasts (coverage) Independent validation, stress scenarios, kill criteria before capital is committed
Change-control / systems failure Knight Capital’s 2012 trading software incident — material loss and near-failure after a bad release (SEC materials) — still the classic MRM teaching case Pre-production testing, release controls, kill switches, independent challenge
Cyber attack surface Prompt injection, over-privileged agents, poisoned context, AI-accelerated social engineering Identity controls, tool permissions, red-team tests, incident playbooks that include “disable the agent fleet”
Shadow AI & IP loss Unmonitored tools become the real operating model; know-how walks into a vendor log Discovery scans, approved alternatives, attestation
Bias / unfair outcomes Credit, pricing or servicing models that treat customers unevenly under Consumer Duty Fairness testing, outcome MI, challenge by Line 2
Correlated / concentration failure Many teams on one model or one cloud LLM — one outage or one bad update hits several important business services Resilience mapping, exit plans, diversification where needed
Accountability vacuum “The model did it” meets SM&CR — no SMF can show reasonable steps Named owners, change logs, reconstruction packs

Regulators will not invent a new excuse for you. They will ask who owned it, what it was allowed to do, what it did, how you spotted it, and what you changed.

Are you in control?

Answer these without a slide deck:

  1. Can you list every live AI use case — including the unofficial ones — with a named owner today?
  2. For each material case, can you show the last independent review date and the open issues?
  3. If the vendor changed the model yesterday, how would you know?
  4. Can you disable it in under an hour without taking an important business service down?
  5. If a customer complained about an AI-influenced outcome last Tuesday, could you reconstruct the decision path by Friday?

If any answer is “not really”, you are not in control. You are in production.

What good looks like right now

Good in 2026 is boring on purpose:

What good looks like
A living inventory with risk tiers and SMF / business owners
Human gates written into workflows, not into posters
Monitoring that produces exceptions humans actually clear
Line 3 coverage in the audit plan, themed across AI — not a one-off project review
Vendor AI treated as outsourcing / third-party risk where it is material
Board / audit committee MI that shows inventory growth, incidents, overdue validations and kill-switch tests — not “number of AI ideas”
Evidence packs that a reasonable SMF would be willing to sit behind
Optional structure from NIST AI RMF or ISO/IEC 42001 — useful scaffolding, not a substitute for FCA/PRA outcomes

That is enough to scale the uses that deserve scale — and to stop the ones that do not.

How to get there: a steady-state path

<!–CHEVRON–>

Stage Outcome
1. Inventory & classify Complete list of AI uses, data classes, autonomy, customer impact
2. Risk tier & owners Every material case has a business owner and a challenger
3. Controls & testing Gates, logging, evaluation harness, change control designed in
4. 1LoD / 2LoD assurance Operating sampling + independent challenge with open-issue tracking
5. Independent validation External or truly independent review where impact warrants it
6. Continuous monitor Drift, incidents, vendor changes, outcome MI on a fixed rhythm
7. Steady-state audit cycle Line 3 plan, board reporting, refresh on triggers — then repeat

Skip stages and you get theatre. Finish stage 7 and AI stops being a special project. It becomes another controlled capability.

How GRT Consulting can help

GRT works with operational, risk, compliance and internal audit teams who need AI that survives contact with FCA expectations — not another principles PDF.

We help firms:

How we help
Shape an AI audit and assurance strategy that fits your permissions, not a generic tech maturity model
Build the inventory, tiering and owner map so SM&CR lines up with reality
Design 1LoD / 2LoD operating rhythms and Line 3 scopes that cover generative and agentic use cases
Stand up practical controls — human gates, observability, evaluation, kill switches, third-party overlays
Bring independent challenge on material models and vendor AI where you need an outside view
Train teams to supervise agents so productivity gains do not outrun control

If you want a structured conversation about whether you are actually in control of your AI — and what to fix first — talk to us.

T: +44 20 3695 9251 E: info@grtconsult.com Web: grtconsult.com


Sources

.., 17th September 2026

GRT Consulting

Speak to us about how we can help you

T: +44 20 3695 9251

E: info@grtconsult.com

Submit Request for Proposal