How do we know if our AI systems are safe and reliable? This summer, Ofgem consulted on AI assurance.
I started my career as a Chemical Engineer – the safety culture never leaves you. And the ALARP principle (As low as reasonably practicable) is the foundation of my response. The consultation closed on 12 August. Link to my full response is at the end. The short version…
We have been assuring AI for conformance and calling it safety.
We publish the principles. We complete the assessments. We certify the management system. None of it evidences the thing that matters: whether the system delivers a safe, fair outcome in the live grid and the real customer decision. A model can tick every box on a governance checklist and still get the decision wrong where it counts. The regulator will not be satisfied by a framework. It will want evidence.
AI does not fail the way our risk registers expect.
Traditional software failed loudly: it crashed, the report came back empty. AI fails quietly. It agrees. The output stays confident and fluent while the information behind it has gone stale or wrong, with nothing on the surface to tell you which. Our assurance was built for systems that break. This one fails by sounding plausible.
You don’t have a vendor moat. It is a siege.
Bespoke, single-vendor AI that is expensive, hard to replace, and captures the operator a little more with every release. The sector’s own project records show the risk arriving in two shapes.
Breadth: the same small pool of vendors recurring across licensee boundaries, so one flawed model assumption becomes a correlated, sector-wide failure. Reliability engineers call this common-mode failure.
Depth: a single licensee returning to the same vendor until it loses the practical ability to challenge or replace what it has bought, and assurance collapses into taking the vendor’s word.
An assurance framework designed against one shape will miss the other. The fix belongs in how innovation funding is designed: fund platforms where models can be swapped, and build the common sector data model the UK has so far stopped short of. Offshore oil and gas did this with OSDU, and it accelerated everything built on top.
The sector does not need a new discipline for any of this. It already owns one.
It is called safety management. We know how to understand a hazard, put proportionate mitigation in place, and provide proof. The organising principle is the one every duty-holder already lives by: risk as low as reasonably practicable. AI is a new kind of hazard, not a reason to start again.
Two additions to Ofgem’s seven principles.
Make independence a principle: assurance that is independent of the vendor who built the system and of whoever set the strategy, the way audit has always worked.
And ask the killer question: should we even use AI? Where a reliable deterministic option exists, that is the first choice, not AI. As I always say, AI’s an ingredient, not a meal.
And if this kind of scrutiny makes you uncomfortable, then that’s exactly why you should stop trusting your vendors and start demanding independent verification.
I am running a small number of independent reads this autumn for energy businesses that will have to answer for their AI under RIIO-3. It starts with a 15-minute call where you tell me about one AI initiative and I tell you what I’d do in your shoes. .