AI Systems in Healthcare Must Be Assessed Like Doctors, Commission Advises
- Jun 23
- 3 min read

Artificial intelligence systems deployed within the NHS and the wider healthcare sector should be regulated using the same staged progression of trust applied to human clinicians, according to a new report from the National Commission into the Regulation of AI in Healthcare. The recommendation marks one of the most direct attempts yet to translate the established framework of medical training into rules governing machine decision-making.
The commission's central argument rests on a simple comparison. A newly qualified doctor does not perform unsupervised surgery on their first day. They progress through years of supervised practice, sitting exams, working under consultants, and gradually taking on greater responsibility as their competence is demonstrated and recorded. The commission wants AI systems operating in clinical settings to be held to an equivalent standard, with their capabilities tested and verified before they are trusted with decisions that carry real consequences for patients.
The technology under scrutiny is what the report terms "agentic AI". This refers to systems capable of planning and carrying out tasks independently, with little or no human checking each step along the way. Unlike simpler tools that flag an abnormal scan for a radiologist to review, agentic systems might triage patients, adjust treatment plans, or coordinate parts of a care pathway largely on their own initiative. The commission was clear that such systems should not be handed high-stakes responsibilities from the outset. Much as a junior doctor must earn the right to operate with greater independence, an AI system would need to show, over a sustained period and across a recorded body of cases, that it performs reliably and safely before being permitted near higher-risk procedures or decisions.
To put this principle into practice, the commission has proposed what it calls a tiered regulatory framework. Under this model, the degree of regulatory scrutiny and the strength of accompanying risk controls would be set according to how much autonomy a given system has been granted. A tool that simply offers a second opinion to a human clinician, who retains final say, would sit at a lower tier and face lighter oversight. A system permitted to act and intervene without a clinician's sign off would sit considerably higher, triggering more rigorous testing, monitoring and audit requirements. The logic is proportionate rather than blanket: regulation tightens in step with the independence the system is given, rather than applying uniform rules to every form of AI regardless of what it actually does.
The proposals emerged from a dedicated session within the commission's broader inquiry into how regulators should respond to the pace at which autonomous medical technology is developing. Officials and clinical advisers involved in the discussion were said to be concerned that existing regulatory tools, largely built around software as a fixed, predictable product, are poorly suited to systems that learn, adapt or act with growing independence over time.
Commissioners were broadly supportive of building the new framework around the principle of autonomy, but stopped short of presenting it as complete. They acknowledged that the level of independence an AI system has is not the only factor that ought to determine how tightly it is governed. The report calls for further work to establish what other risk factors, such as the clinical area in which a system operates, the population it serves, or the consequences of failure, should also shape the rules. That investigation is expected to inform the next phase of the commission's work, with a fuller set of recommendations anticipated in due course.
For now, the message from the commission is that trust in medical AI cannot simply be assumed. It must be built, tested and proven in stages, in much the same way it always has been for the people who work alongside these systems.



