Ashley Hirst Writing on community, artificial intelligence and insurance
Insurance & risk

Learning by experience in a world of AI

If AI takes over routine underwriting, the answer is not to preserve a hand-underwriting quota: it is to route exceptions to humans and make replay the training ground.

Ashley Hirst · 24 August 2026 · 5 minute read

There’s a discussion that keeps coming up at least in the (somewhat obscure) world I work in. If AI is ‘taking over’ more of the routine work, how will a new generation of talent learn the required skills to do the advanced work? In a previous article I tackled the general problem of the ‘deal’ that new talent represents to companies changing in relative attractiveness. In this one, I want to tackle how we can reasonably set up new talent for success in learning skills that are required for a high-skilled, knowledge-based role.

In the absence of obvious ‘good’ answers, one suggestion I keep hearing is that talent should keep doing a percentage of work by hand, so the next generation still learns the craft. It sounds prudent. It is the kind of idea that survives a steering committee because nobody can object to “preserving judgment.”

I think it is a category error, and the history of the industry’s own tooling shows why.

When models moved from paper to Excel, nobody kept doing 10% of the calculations by hand to stay sharp. They trusted the spreadsheet, and mental arithmetic stopped being a professional skill anyone mourned. When actuaries took over model building, underwriters did not build their own models 10% of the time. The modelling work moved to the people best placed to do it, underwriters kept the parts of the job that were genuinely theirs, and the freed time went into distribution, broker relationships, and portfolio work. In neither case did anyone preserve the old activity as a training ritual, and in neither case did judgment die. It moved.

Why the 10% fails even as pedagogy

Set aside the economics and take the proposal on its own terms: does hand-underwriting a random 10% of cases actually teach anything?

Mostly no, and the reason is which tenth you get. A random sample is dominated by the median case, and the median case is exactly where the model is most reliable and the human adds least. Nobody’s underwriting judgment was ever built on median cases. It was built on the weird submission, the broker who oversold a risk, the wording that did not quite fit the peril, the loss that came from a direction the file never contemplated. The proposal preserves the least instructive slice of the old job and calls it a curriculum. If the old role was 90% routine cases teaching you very little per hour, mandating a sample of that routine is preserving the boredom, not the learning.

A worked proof from current practice

Insurance already contains a proof that case-level handwork is not a necessary condition for deep underwriting judgment: treaty reinsurance.

A treaty underwriter never sees the individual risks. They underwrite the cedent: the quality of the data, the incentives of the people supplying it, the drift in the portfolio, actual-versus-expected, rate change monitoring, the diagnostics that tell you whether a book is what it claims to be. That is a deep form of judgment, built over years, and it is built entirely one level up from the case.

Follow that logic forward and the endpoint of AI-driven underwriting is not “underwriters plus a model.” It is that the person who oversees the model, today probably an actuary, becomes a portfolio underwriter with enormous productivity relative to the case-level role, and the case underwriter as we knew it disappears. There is no residual value in the 10% task, just as there was no residual value in checking Excel’s arithmetic. This gets uncomfortable for a lot of people whose identity is bound up in the case work, and I think the industry will spend several years pretending otherwise, but the treaty precedent says the judgment survives the move.

Sounds awful, and it comes with a catch

Treaty underwriting works because someone downstream still touches the cases. The cedent’s team does the primary underwriting, and that activity generates the information the treaty underwriter reads. If AI eats the case layer entirely, the portfolio underwriter is monitoring a system whose contact with ground truth is itself model-mediated, all the way down.

Aviation ran this experiment first. Cockpit automation did not fail pilots on routine flights; routine flying is precisely what it does better than humans. It failed them in the handback, the moment the automation gave up a degraded aircraft to a crew whose instincts for the exception had atrophied - Air France 447 is the case every safety course teaches. The insurance version of that failure is not mispricing the book on average, because model diagnostics will catch that. It is the novel exposure accumulating silently, the next silent cyber, the wording that responds to a peril nobody contemplated, because those things present first as themes in individual cases, long before they present as a signal in portfolio data. If no human is reading cases, nobody sees the themes.

Exceptions not quotas

So the redesign I would argue for is not a handwork quota. It is an exception surface. The model routes to humans exactly what it is least confident about: the off-distribution submission, the wording it cannot parse, the segment that is drifting. Humans work there and only there.

This is audit logic rather than apprenticeship logic, and I think it is a genuinely better curriculum rather than just making an arbitrary compromise. Investigating anomalies is how claims-side and forensic people have always built distinctive judgment, and it is denser learning per hour than a caseload of medians ever was. Add the capability AI makes nearly free, retrospective replay: re-underwrite last year’s book against actual outcomes, at will, a hundred times if you like. A junior with an exception queue and unlimited replay may plausibly accumulate more real judgment in three years than the old apprenticeship delivered in ten. That is the compression argument I made in my previous article.

One quite subtle problem remains genuinely unsolved. Recognising an anomaly assumes talent knows what normal looks like, and today that prior comes from having done the median work the AI now does. The transition generation is covered; they built their priors the old way. The steady state, where someone builds the prior from exceptions and simulation alone, is an experiment nobody has run. Anyone who claims to know it works is guessing, and so is anyone who claims it can’t.

My prediction: the 10%-by-hand proposals will get adopted in several places, decay into a compliance ritual that nobody learns from, and be dropped within a few years. The firms that get this right will be the ones that build exception routing and replay into the workflow from the start and treat the anomaly queue as their training ground. I could be wrong about how fast the ritual version dies. I do not think I am wrong that it is a ritual.

As usual, I look forward to being told where I’m wrong.


Portrait of Ashley Hirst

I work in insurance and write about artificial intelligence, risk and community — Jewish and British. This site collects the writing. More about me.