A no entry sign over an image of a reverse-centaur where the robot is in control

Last month the Institute and Faculty of Actuaries (IFoA) launched what it calls its “AI Manifesto“. I thought I would take a look.

The Manifesto is admirably short and to the point. Its pitch can be summarised by this final line of its Introduction and context:

For the benefit of society, we encourage those responsible for developing and deploying AI systems to ensure that there is an actuary in the loop.

Think of it as a massive “Gissa job” plea to those companies currently investing in AI systems. But what sort of jobs are we pleading for?

The idea is that actuaries have been working with models which use machine learning for many years, that we have statistical expertise, are comfortable working with data, used to stress-testing systems and also to looking at the risk appetite of a business as a whole.

I think the first hint that there might be a problem here is when the Manifesto moves onto ethics and the public interest. As it says:

Actuaries help teams balance commercial aims with societal expectations, offering constructive solutions rather than imposing limits.

It is difficult to square this with Anthropic‘s recent IPO filing, where they warn its technology may pose “existential risks to humanity”. I am pretty confident that Society expects not to be put at existential risk. So if Anthropic thinks its technology is that dangerous, there is no balance to be found between commercial aims and societal expectations here. Societal expectations are paramount. And limits will need to be imposed rather than “constructive solutions”.

But, in my view, the main problem here is what “real world” the IFoA feel they are serving. This comes through in the role for “professionalism” they envisage:

Projects with sound validation, honest documentation and clear accountability clear governance more quickly, avoid late-stage rework, and fail in testing rather than in production. Professionalism is what lets an organisation say “yes” to AI with confidence – and protects it from the adverse outcomes that would otherwise force it to say “no”.

This, remember, is in an environment where the latest large language models like Anthropic’s Claude Opus 5.5 or OpenAI’s GPT-6 Astra, are proprietary models where validation, documentation, accountability and governance are all just what they let you see. A framework which allows companies to say yes with misplaced confidence is unlikely to be serving the public interest. And translating a capacity to apply a regulatory model to an insurance company where you have all the data and can make informed judgements about the things that you don’t know (Actuaries are used to working with large, imperfect, transactional datasets subject to audit and regulatory scrutiny. We clarify what data is needed, what quality level is acceptable, and how data should be checked, reconciled and monitored over time) is emphatically not the same as working with a model where you don’t know precisely how it works, cannot accurately predict the errors it will make and do not have enough time to look in all of the places you need to.

Take bias for instance. The Manifesto says:

We examine how a model behaves for different groups and sub-groups, test whether training data is representative of the people the system will affect, and help manage the biases we find responsibly through fairness metrics for predictive models, and output audits and red-teaming for generative systems.

Nita Farahany in the law and policy course documented in her Thinking Freely Substack demonstrates the limits of that bias management in practice, and this is Cory Doctorow on red-teaming (from Red Team Blues):

That’s the problem with blue teaming it—you need to be perfect, while—The red team only has to find a single error

Yes, that’s right, the IFoA haven’t even understood that they are on the Blue Team when it comes to AI errors, ie the one supposed to be stopping as many of them as possible. The Red Team breaks into systems by exploiting a single vulnerability. That is comparatively easy. The Blue Team needs to reassure its clients that it has removed ALL vulnerabilities.

That will be the role of any “actuary in the loop”. And the chances are it won’t be possible. The three examples given by the IFoA are as follows:

A generative AI customer assistant
An insurer wants an LLM-based assistant to answer customer queries. An actuary designs the evaluation
framework: a curated test set of realistic queries, explicit tolerances for error, escalation rules for high-stakes topics, and live monitoring of answer quality. The project clears internal governance at the first attempt because approvers can see exactly what “good enough” means and how it is evidenced.

A machine-learning decision model
A lender deploys an ML model to score applications. An actuary tests behaviour across customer sub-groups,
identifies a data artefact that disadvantages one segment, and works with the data science team to correct it. The documentation of that judgement becomes the centrepiece of a constructive conversation with the regulator.

AI risk across an organisation
A Chief Risk Officer faces AI adoption in a dozen business units at once. An actuary helps articulate an enterprise AI risk appetite, maps where the organisation depends on the same foundation models, and builds board reporting that distinguishes experimentation from deployment. The organisation accelerates adoption in low-risk areas precisely because it now knows where the high-risk areas are.

None of these examples look unreasonable in isolation, but, to justify the eye-watering expense of this technology, it will need to increase speed and capacity significantly. So, rather than a single generative AI customer assistant, imagine a client relationship management role where a junior actuary or senior student manages 10 times the current expected number of clients at the same time, because the generative AI customer assistant makes that possible?

Will the skills of an actuary versed in the intricacies of Solvency 2 or pensions regulation be so revered by your client that you would be given more than enough time to rubber stamp a system serving 30 or 40 clients? How long do you think you’d need to check the output on all of them properly? I think it unlikely you will be given that long.

This is the risk of becoming a reverse centaur, where you work for the assistant rather than the other way around. And my view is that the actuarial education system is positively encouraging the development of reverse centaurs amongst our current students.

It gets worse. The early evidence suggests that working with AI systems encourages rubber stamping. In one study of AI-generated audits with an auditor in the loop, 150 of the 160 reports were submitted without any changes made by the human. The performance of humans in the loop may also be adversely affected by working with AI systems, leading to “agency decay”:

Agency decay is the gradual erosion of a person’s ability and willingness to observe carefully, think independently, choose deliberately, and act responsibly. It does not arise because AI is inherently harmful. It arises when convenience becomes the default setting for cognition. A tool that first helps us think can begin to think around us, then for us, then without us noticing what has been weakened.

There are lessons to be learned from other professions who have been working for much longer with expert systems. The aviation industry for example. As Craig Bright, Co-Chief Operating Officer at Barclays has posted:

So the test shouldn’t be whether a human appears somewhere on the process map.

It’s whether the system remains understandable, controllable and recoverable when the agent is wrong, uncertain or unavailable.

Safety doesn’t come from leaving a human in the loop.

It comes from designing a system the human can still control.

The IFoA’s position appears contradictory. On the one hand, they clearly don’t believe that AI is as powerful or as dangerous as Anthropic are claiming in order to push up the IPO share price, or I assume they would not be pushing so hard to get an actuary on every team “developing and deploying AI systems”. I think they are right not to believe it, for a number of reasons, but Naomi Alderman’s AI predictions: seven ways you can tell useful thinking from sci-fi fantasies gives probably the most entertaining ones.

On the other hand, if this is just another technology for which the price is going to need to adjust to reality at some point, we need to be very careful as a profession not to be part of the mob urging their clients to adopt it at a faster rate than they would otherwise. Particularly if it is damaging our own professionals in the process.

AI in some form will survive the inevitable crash, but the shape and scope of the wreckage that will be left behind is currently hard to predict. Let’s be careful what we wish for.