
An artificial intelligence transcription tool used by the NHS in England told a patient they had demyelination, the nerve damage behind conditions such as multiple sclerosis. The original test result had read “null demyelination” — meaning no demyelination was found. The AI had simply dropped the word that reversed the meaning.
That case is among the examples collected by Healthwatch England, the statutory patient watchdog, and reported widely today. The finding is that AI scribes now transcribing consultations for GPs and hospital doctors are regularly getting drug names and diagnoses wrong. In many cases, the patient is the one who spots the mistake.
Key facts at a glance
- 27 different AI scribes are in use across the health service in England.
- One scribe swapped a prescribed drug for another with a similar name.
- Another omitted a consultant’s instruction to seek a repeat prescription for migraine medication.
- A third recorded a doctor telling a patient to continue taking Prozac when the doctor had neither prescribed nor discussed it.
- AI scribes are not classified as medical devices by the Medicines and Healthcare products Regulatory Agency.
What is an AI scribe?
AI scribes are ambient documentation tools. They sit in the consulting room, listen to the conversation between clinician and patient, and automatically produce a clinical note. That note is entered into the patient’s permanent record, and often a version is sent to the patient as a letter.
The pitch is simple: doctors spend hours every day typing notes and letters. A tool that can do this in real time frees clinicians to focus on the patient. It is not hard to see why the tools have spread quickly across NHS trusts and GP surgeries.
But the spread has happened faster than the safety checks. Healthwatch England’s warning makes clear that the systems produce fluent, plausible notes even when they are wrong. Because the output looks normal, it can survive the busy glance a clinician gives before signing it off.
The high cost of small errors
The errors collected by the watchdog are mundane in form and serious in effect. One scribe swapped a prescribed drug for a different drug with a similar name — the kind of confusion pharmacology spends considerable effort designing out. Another summary omitted a consultant’s instruction about a repeat prescription for migraine medication. A third recorded a doctor telling a patient to continue their Prozac, when that doctor had neither prescribed nor discussed it.
These are not the spectacular hallucinations that dominate AI headlines. Nobody was misled by a fabricated statistic or an invented source. A real sentence was rendered slightly wrong — but slightly wrong is sufficient when the sentence names a drug or a diagnosis.
Demyelination is a case in point. The phrase “null demyelination” is a negative result. Dropping the word “null” turns a reassuring finding into a life-changing diagnosis. If the patient had not been told to check, or had not noticed, the false positive could have remained in their record permanently.
No regulatory oversight
The deeper problem is structural. There is no England-wide oversight of these tools. The Medicines and Healthcare products Regulatory Agency has not classified AI scribes as medical devices, which places them outside the regime that would test them for safety and effectiveness before deployment.
The regulator did publish guidance in August clarifying where the line sits. A system that only transcribes what was said is not a device. A system that suggests a diagnosis or a treatment may well be a device. That places a great deal of weight on how each product is described by its vendor.
The incentive that follows is obvious: a scribe marketed as a passive transcriber avoids a regulatory process that a scribe marketed as a clinical assistant would have to complete. The same underlying model can be described in different ways, and the description determines whether it falls under the medical device regime.
Why the mistakes go unnoticed
Part of the problem is that the errors are invisible to the people who should be catching them. A busy clinician may not read every word of an automatically generated note. The patient, meanwhile, is the least equipped person to know what the note should have said, yet they are often the last line of defence.
Healthwatch England warned: “These inaccuracies may persist in their records if the patient doesn’t catch them.” That puts the responsibility on the person with the least information.
Rachel Power, chief executive of the Patients Association, is among those raising concerns. She is joined by clinicians including London GP Dr Shier Ziser Dawood and Charlotte Blease of Uppsala University in Sweden. The objection is not to the technology itself, but to its arrival without a safety net.
The trade-off that drives adoption
None of this explains why these tools spread so quickly. Clinical documentation is the administrative burden doctors complain about most. A system that reliably removes an hour of typing a day will be adopted whether or not anyone has assessed it.
The problem is what happens in between. A tool adopted for its speed, unassessed because of how it is categorised, producing a document that becomes the permanent clinical record — this is a chain in which no single link is obviously anyone’s responsibility.
The result is a system with 27 different products, no device classification, and a check performed by whoever happens to read their letter carefully.
What should happen next
The remedy Healthwatch England points to is unglamorous and probably right. Patients should be told when a scribe is being used and given their notes to check. That turns an accidental safety mechanism into a deliberate one.
But patient checking alone is not enough. Clinicians need a reliable way to verify AI transcription before signing it. Trusts need to know which products have been tested. Regulators need to close the classification gap so that a vendor cannot dodge scrutiny by calling a clinical assistant a transcriber.
The wider lesson is that fluency is not accuracy. A machine that generates confident English sentences can be wrong in ways that are difficult to detect. When those sentences become part of a medical record, the consequences are not theoretical.
AI scribes may well reduce burnout and improve the experience of consultations. But the technology has been introduced before the system that should govern it. Until there is oversight, transparency, and a way for mistakes to be caught and corrected, patients in England are carrying the risk.
