Whose job is it to notice?

Sonya Cullington is a cyberpsychologist and digital policy advisor. She works on what sustained AI use does to human cognition, professional identity, and judgement, and what organisations and individuals can do about it. She is the creator of the AI Judgement Framework, the founding architect of the NHS Communications AI Taskforce, and chair of the Patient and Public Advocacy Steering Committee at UK Digital Health and Care.

The Health Foundation's deliberative research for the National Commission into the Regulation of AI in Healthcare tested a skin cancer detection tool with members of the public. In clinical studies it performed slightly better than a human dermatologist. Participants were told it could clear around 30 per cent of patients with negative results without any human review. In all three workshop locations, they rejected that. People who supported AI in healthcare, and who had seen the accuracy data, still wanted someone with clinical responsibility to look at the result.

That answer ran through the Commission's whole evidence programme. More than 12,000 people contributed, including members of the public, 31 professional bodies, patient groups and a regulatory sandbox testing seven AI products in real-world conditions. Asked what would make AI in healthcare safe and trustworthy, they kept returning to the human in the system. Seventy per cent of the public prioritised human oversight over speed of results.

The professional bodies explained why that oversight is fragile. They named automation bias under time pressure, the difficulty of challenging an AI output in a culture that does not support it, particularly for less experienced staff, and de-skilling as long-term reliance wears away independent clinical reasoning.

The AI Airlock sandbox saw this happen. It tested TORTUS, an ambient voice tool for clinical notetaking, which performed well, with a precision of 0.989. As staff got used to output that was almost always right, they checked the detail less closely, and the errors it still made became harder to catch. Governance frameworks treat human oversight as a constant. The Airlock suggests it behaves more like a dial, and familiarity turns it down.

The Commission's 44 recommendations build a strong framework for the technology. Tools will prove themselves in supervised real-world settings before full approval, oversight continues across a product's life, and liability is shared so clinicians are not left carrying all the risk. Where the recommendations reach the people using these tools, the main lever is training. Training happens at the start and is aimed at individuals. The problems the evidence describes build over years and come from how the system operates. A clinician can be well trained on day one and check less carefully two years later.

Read the evidence and the recommendations together and five human risks emerge that nobody has been asked to own.

The first is independent reasoning. The professional regulators are named as training partners, with a brief focused on competence to use AI. No body is asked to track whether clinical reasoning holds up after years of working alongside it.

The second is the ability to say no. A resident doctor who senses something is wrong with an AI output needs to know the consultant will back them and the trust will treat it as good practice. No recommendation says who creates that culture, or whether the Care Quality Commission (CQC) will inspect for it.

The third is the relationship between patient and clinician. R25 asks CQC to work with patients on what good care looks like once AI is involved. That is a real step, and there is still nothing to measure the relationship against, because patient experience checks and professional standards were written before AI was in the room.

The fourth is the definition of harm. The system is built to catch faulty devices. A clinician whose reasoning has dulled, a nurse who has stopped speaking up, a patient whose appointments feel more like being processed: in each case the tool is working as designed, so nothing gets reported.

The fifth is the cost of checking. The report itself notes that many clinicians feel "overly burdened with the responsibility of monitoring for drift". Every tool that needs a human check adds to that load, and nobody in workforce planning is counting it.

R30's proposed national observatory could hold several of these. As drafted, it would track organisational readiness and tool performance. If it also tracked what sustained AI use does to people, the system would have someone whose job it is to notice.

The government's formal response is still to come. Until it names owners, the framework protects the technology and leaves the humans to look after themselves. Organisations deploying AI now can start answering these questions themselves:

  • Who in your organisation is responsible for noticing whether staff check AI outputs as carefully in year two as they did in month one?

  • If a junior member of staff overrides an AI output and turns out to be right, would anyone know, and would it be recognised?

  • Is the time staff spend checking AI outputs counted anywhere in your workforce plan?

National Commission into the Regulation of AI in Healthcare: Recommendations for a future regulatory framework: https://www.gov.uk/government/publications/national-commission-into-the-regulation-of-ai-in-healthcare-recommendations-for-a-future-regulatory-framework

To find out more, visit sonyacullingtonconsulting.com or connect with Sonya on LinkedIn.

Next
Next

Power over Ethernet times Endpoints equals Smart, Sustainable Networks