• Insights · AI Governance
Case Study: AI-Assisted Underwriting — Workflow Redesign in a Financial Services Back Office
In 2026, the AI pilot became the easy part. The hard part is proving how every AI-touched underwriting decision was made — and that is what regulators will care about in 2027.
By Suchetana Bauri · 16 September 2026 · AI Strategy & Governance
The Challenge
AI can prepare the underwriting case. But if the workflow does not expose evidence, conflicts and authority, it smuggles decisions past the people responsible for them.
The Opportunity
Redesign work so the assistant cannot hide uncertainty, erase provenance or become the unnamed author of the decision.
Until recently, the underwriter’s job was to find those facts by hand. Now an AI assistant can extract them, compare them with appetite rules and draft a referral note in minutes. The obvious sales pitch is speed. The more important question is what happens to judgement when the machine prepares the case.
Why 2026 changes the stakes
That question is no longer theoretical. In February 2026, EIOPA reported that nearly two-thirds of 347 surveyed European insurers were already using generative AI. Most remained at proof-of-concept stage; however, 64% of reported use cases targeted back-office productivity.
The direction is clear. Insurance is moving from scattered experiments to production systems. The real challenge is not whether AI can summarise a risk file. It is whether an insurer can redesign the work around that summary without weakening evidence, fairness or accountability.
This composite case study follows one commercial insurer through that redesign. The company and characters are fictional, but the workflow, pressures and controls reflect current industry practice.
01 –
The case for change
Northstar Commercial Insurance, a composite mid-sized UK insurer, writes property and liability cover through brokers. Its underwriters handle established products, but the submissions are rarely standard.
The company’s service promise says it will acknowledge new business within four working hours. In practice, complex cases sit in queues while operations staff rename attachments, copy values into the policy system and ask for missing documents.
The process looks digital because the documents arrive electronically. It is not. It is manual work performed on screens.
The 2024 Bank of England and FCA survey found that 75% of responding financial firms already used AI, while another 10% planned to do so within three years. Yet 62% of reported use cases were considered low-materiality, against only 16% at higher materiality.
AI Adoption
75%
of financial firms actively use AI
Complexity
62%
cases deemed low-materiality
High Materiality
16%
operating at critical depth
02 –
Shadow AI was already there
The formal project began after a data-loss-prevention alert. An operations analyst had copied sections of a broker email into a public generative AI service to produce a summary. No customer name appeared in the prompt. Nevertheless, the text included an address, turnover, claims details and the name of a director.
The analyst was not trying to bypass security. She was trying to clear a queue.
This is the central point: employees reach for unsanctioned tools because the sanctioned workflow gives them too much friction and too little help. A warning banner may deter some use. However, it cannot make a slow process fast.
Northstar’s first instinct was to block more websites. Its second instinct was better: examine the job the analyst had been trying to do. She needed to identify entities, summarise documents, compare conflicting values and list missing information. Those were legitimate tasks. The failure lay in providing no safe way to complete them.
The programme did not begin with a model. It began with a workflow map.
03 –
The old workflow
Legacy Operational Pipeline
01.
Download
Operations analyst pulls multi-format broker files
02.
Classify
Manual document sorting & basic detail transcription
03.
Review
Underwriter parses files individually, cross-referencing rules
04.
Clarify
Manual lookup and query generation for broker gaps
05.
Escalate
Drafting formal peer/committee referral sheets
06.
Record
Manual transcription of final parameters into policy ledger
Each stage created its own notes — in workflow tools, email, spreadsheets or personal folders. The policy system recorded the final answer but not always the evidence trail.
The team measured turnaround time from receipt to quote. That encouraged queue clearing, not better decisions.

04 –
The redesigned flow
AI-Assisted Workbench Operations
1. Automatic document intake, filtering duplicate/unreadable attachments
2. Entity fact mapping linked directly to original source documents
3. Conflict detection flags discrepancies in financials across files
4. Verification checks against active corporate risk criteria
5. Automated first-draft risk memo compilation for reviewer approval
The evidence link changes the nature of the automation. Instead of asking the underwriter to trust a fluent paragraph, the system asks them to verify a series of claims.
If two documents disagree, the workbench displays both values rather than choosing one silently. If no evidence supports a generated statement, the statement cannot enter the formal case summary.
The interface is slower than a blank chatbot and faster than manual review. That is the right trade-off. In underwriting, friction is not always waste. Sometimes it is the moment in which a person notices that the machine is wrong.
05 –
Human oversight means work
“Human in the loop” has become a decorative phrase. It can mean anything from careful review to a tired employee clicking a green button.
Northstar replaced the phrase with specific obligations: verifying material extracted facts, resolving conflicts, recording the reason for any override, writing the final rationale, and requiring a second reviewer for high-risk cases.
Crucially, the system records edits. If the AI drafts a referral saying a building has no combustible cladding and the underwriter changes it, that correction becomes monitoring data.
A falling override rate does not automatically prove improvement. Underwriters may trust the tool more because it is better. They may also have stopped challenging it.
06 –
Accountability enters the screen
Socio-Technical Governance Map
01.
Underwriting Director — Decision Authority
Accountable for pricing decisions, alignment, and long-term customer outcomes
02.
Operations Lead — Process Quality
Intake quality control, human workflow timelines, queue monitoring
03.
Model-Risk Specialist — Validation & Oversight
Assuring prompt safety, extraction compliance, drift alerts
04.
Technical Owner — Platform Resilience
Identity tokens, pipeline maintenance, system change logs
01.
Procurement Manager — Contractual Alignment
Vendor escalation protocols, regulatory support evidence
One named executive remains accountable for the use case. The governance forum can challenge that person, but it cannot absorb the responsibility.
Each rule has an owner, while model versions carry approval records. Material alerts follow a defined escalation path, and every override remains traceable to a named underwriter.
07 –
The 2026 regulatory turn
The most important regulatory shift this year is from principles to proof.
The FCA’s AI Live Testing programme assesses the complete deployed system: its context, governance, human controls, evaluation and input-output safeguards. Its second cohort began testing in 2026.
The government’s July 2026 Financial Services AI Adoption Plan recommended voluntary sharing of AI incidents and near misses and proposed an industry-led assurance framework for third-party AI providers.
Northstar treats unsupported recommendations, missed referrals and inappropriate data disclosures as reportable internal events, even when no customer suffers measurable harm.

08 –
The model is not the system
Hierarchical Assurance Validation
Level 1: Component Validation
Factual extraction & syntax matching. Checks if model reads and categorizes fields accurately.
Level 2: Workspace Routing Integration
Ensuring proper data flows. Do warnings reach the designated underwriters securely?
Level 3: Underwriting Portfolio Audits
Aggregate output patterns. Evaluating variance, decline percentages and pricing fairness.
The company also maintains a versioned inventory recording: model, supplier, purpose, permitted data, affected products, decision influence, owner, limitations, validation status and fallback procedure.
As of April 2026, 25 US jurisdictions had adopted the NAIC model bulletin and 12 states were piloting an AI Systems Evaluation Tool.
09 –
Fairness cannot wait
Underwriting necessarily distinguishes between risks. Fairness does not require identical prices. However, it does require understanding which data drive differences and whether proxies create unjustified effects.
Northstar tests the assistant before release and after material changes. It compares extraction errors across document types, broker channels and customer groups.
The team pays particular attention to missingness. “Unknown” remains a visible state; the system cannot quietly convert it into “no”.
10 –
Drift is an underwriting problem
AI failures are often imagined as spectacular hallucinations. More commonly, performance deteriorates quietly.
Northstar monitors: unsupported claims, extraction corrections, referral rates, override reasons, processing time, complaint themes and differences between predicted and realised loss experience. It also looks for concentration: a tool that nudges many underwriters in the same direction can create a portfolio problem.
The company maintains a fallback route. If the assistant fails, the team can return to manual processing without losing original documents or decision authority.
11 –
The supplier does not own the risk
The contract covers data location, retention, subcontractors, model changes, incident notification, audit rights and exit support. More unusually, it also covers evaluation evidence.
Procurement initially resisted. Detailed conditions might slow the deal. Yet speed at purchase can become dependence after deployment.
12 –
Measuring the redesign
Balanced Operational Scorecard
1. Service Metrics
Turnaround timeline limits, referral queues, transaction speed, age of bottlenecks
2. Quality Metrics
Fact extraction edits, source matching, syntax corrections, missing field overrides
3. Judgement Metrics
Authority exceptions, secondary peer reviews, risk appetite departures
4. Portfolio Outcomes
Authority exceptions, secondary peer reviews, risk appetite departures
No single measure decides whether the system succeeds. Faster processing alongside rising corrections is not success. The governance group reads measures together.
13 –
What Northstar refused
01
No Automated Declines
Every rejection must be signed off manually by a licensed human underwriting authority.
02
No Generalist Agent
Refused cross-product tools. Bounded task maps for single target domains are safer to govern.
03
No Silent System Updates
Mandated strict version gating. Retested parameters before pushing model upgrades live.
These refusals slowed deployment. They also made deployment possible.
14 –
The 2027 test
Under the revised EU timetable, obligations for Annex III high-risk AI systems apply from 2 December 2027.
By 2027, a credible insurer should be able to reconstruct one AI-assisted decision without launching an internal investigation. It should produce: original submission parameters, specific active model versions, extracted source files, conflict history, human override reasons, and final authorisation logs.
If those answers live across inboxes and personal spreadsheets, the insurer has not built accountable AI.
15 –
The practical lesson
AI-assisted underwriting is often presented as a choice between automation and human expertise. That is the wrong frame.
Northstar did not remove the underwriter. Nor did it preserve every manual task. Instead, it automated retrieval and preparation, exposed uncertainty, strengthened referrals and made responsibility visible.
The bold move in 2026 is not to put an AI assistant beside every underwriter. It is to redesign underwriting so that the assistant cannot hide uncertainty, erase provenance or become the unnamed author of the decision.
In 2027, regulators may ask for the record. Customers may challenge the outcome. The market may expose a drifting assumption. When that happens, “a human was involved” will not be an answer. The answer will be the workflow.
“The answer is not to slow everybody to the speed of governance. It is to rebuild governance at the speed of work.”
