Skip to content

AI in Clinical Practice

Abstract

Medicine has been the favourite demonstration field of artificial intelligence since the 1970s, and its clinics have been where the demonstrations were tested. A Leeds program diagnosed acute abdominal pain more accurately than senior doctors in 1972; Stanford’s MYCIN matched infectious-disease experts on paper and never treated a patient. Computer-aided detection for mammography, approved by the US Food and Drug Administration in 1998 and paid for by Medicare from 2002, was on most American screening mammograms by the 2010s, until a study of 625,000 of them found it improved nothing. Deep learning brought the first autonomous diagnostic device, a diabetic-eye screener authorized in April 2018, and by the end of 2025 the FDA had authorized 1,451 AI-enabled devices, three-quarters of them in radiology. Geoffrey Hinton said in 2016 that hospitals should stop training radiologists; ten years later the United States had more of them, better paid, and not enough. Large language models passed medical licensing questions in 2022. The recurring finding is that accuracy measured on a curated dataset tells little about what happens on a ward.

Leeds, 1972

The first trial that set a computer against doctors on real patients as they arrived was run at the General Infirmary in Leeds by the surgeon F. T. de Dombal and colleagues at the University of Leeds. Their program applied Bayes’ theorem to the recorded frequencies of symptoms in earlier patients and ranked seven diagnoses, the headings under which, they found, more than 95 percent of acute abdominal pain fell. In a prospective study of 304 patients admitted during 1971, published in the British Medical Journal in April 1972, it was right in 91.8 percent of cases; the most senior member of the clinical team to see each patient was right in 79.6 percent. The authors concluded modestly that such a system was feasible “and likely to be of practical value, albeit in a small percentage of cases.” The computer correctly classified 84 of the 85 patients with appendicitis.

At Stanford, Edward Shortliffe’s rule-based MYCIN (1972–76) recommended antibiotics for blood infections and meningitis; in a 1979 evaluation Stanford’s infectious-disease faculty rated its recommendations acceptable in 65 percent of cases, against 42 to 62 percent for practising physicians and students. It was never used on patients; the liability, the time it took to enter a case, and the lack of any place for it in a hospital’s routine stopped it, as they stopped most of the expert systems that followed.

The Mammography Machine

The first AI-like product that most patients actually encountered was computer-aided detection (CAD) for screening mammograms. The software marked regions of an image that looked like masses or calcifications, and the radiologist decided. The FDA approved it in 1998; in 2002 the Centers for Medicare and Medicaid Services began paying extra for it, and it spread to most screening mammograms in the United States at a cost of more than $400 million a year.

In 2015 Constance Lehman and colleagues in the Breast Cancer Surveillance Consortium compared 495,818 digital screening mammograms read with CAD and 129,807 read without it, interpreted by 271 radiologists at 66 facilities between 2003 and 2009. Sensitivity was 85.3 percent with CAD and 87.3 percent without; specificity was 91.6 and 91.4 percent; the cancer detection rate was 4.1 per 1,000 women either way. Among the 107 radiologists who read both ways, sensitivity was significantly lower with CAD. “Computer-aided detection does not improve diagnostic accuracy of mammography,” the paper concluded, and insurers were paying more “with no established benefit to women.” The product had been authorized and reimbursed on the strength of reader studies, not on outcomes in practice.

The Eye Screeners

Fundus Retinopathy NEI
A retinal photograph showing diabetic retinopathy, the kind of image the screening systems grade. Image: National Eye Institute, public domain, via Wikimedia Commons.

Deep learning changed what the software could see. In December 2016 a Google team led by Varun Gulshan reported in JAMA that a convolutional network trained on 128,175 retinal photographs, each graded three to seven times by a panel of 54 US-licensed ophthalmologists and senior ophthalmology residents, detected referable diabetic retinopathy about as well as the graders (see ImageNet and the Deep Learning Revolution). Diabetic eye disease was a good target: people with diabetes need regular screening, and the photograph is the whole input.

In April 2018 the FDA authorized IDx-DR, from a company founded by the University of Iowa ophthalmologist Michael Abràmoff, as the first autonomous AI diagnostic device in any field of medicine: it returned a screening decision without a clinician reading the image. Its pivotal trial had enrolled 900 people with diabetes at primary-care clinics, and it detected more than mild retinopathy with a sensitivity of 87.2 percent and a specificity of 90.7 percent against a reading centre’s grading.

The field tests were less tidy. When Google Health deployed its system in eleven clinics in Thailand, a team led by Emma Beede observed that nurses photographing dozens of patients an hour in rooms that could not be darkened produced images the system refused to grade: of 1,838 images in the first six months, 393, or 21 percent, were rejected, and those patients were told to travel to a hospital for a specialist exam. Slow internet connections delayed the promised instant result. Researchers who tried to reproduce the 2016 JAMA result with public data in 2019 reached an area under the curve of 0.95 on one test set and 0.85 on another, against the 0.99 reported; the original training data and code were not public.

The Sepsis Alarm

Models that run inside the hospital record receive less scrutiny than devices, because many never go to the FDA. The electronic-health-record vendor Epic Systems shipped a proprietary sepsis prediction score that hundreds of hospitals switched on. In 2021 Andrew Wong, Karandeep Singh and colleagues at the University of Michigan tested it on 38,455 hospitalizations. Its area under the curve was 0.63, where 0.5 is chance. It missed 1,709 of the 2,552 patients who developed sepsis, 67 percent, while raising alerts on 18 percent of all patients; of the septic patients whom clinicians had not already treated in time, it flagged 183. The authors wrote that its wide adoption despite poor performance raised “fundamental concerns about sepsis management on a national level.”

Stop Training Radiologists

At a Creative Destruction Lab event in Toronto in 2016, Geoffrey Hinton said: “People should stop training radiologists now. It’s just completely obvious that within five years deep learning is going to do better than radiologists.” Radiology became the specialty with the most AI products: of the 1,451 AI-enabled devices the FDA had authorized by 31 December 2025, 1,104, or 76 percent, were radiology devices. It did not become the specialty with the fewest doctors. By 2024 radiology residents were writing about the largest radiologist shortage in the field’s history; imaging caseloads rose about 25 percent between 2018 and early 2025, the US radiologist workforce grew about 10 percent in a decade, and average pay reached $571,000. In 2023 Hinton moved parity to “another 10 or 15 years,” and in 2025 he said he had meant the reading of images, not the profession. Medicare and Medicaid still require a licensed physician to make the final read.

Language Models

IBM Watson Health had tried from 2011 to turn a question-answering system into an oncologist and was sold off in 2022. Large language models reached medicine by a different route: they were trained on general text and turned out to answer medical questions. In December 2022 a Google team led by Karan Singhal reported that Flan-PaLM scored 67.6 percent on MedQA, a set of US Medical Licensing Examination-style questions, more than 17 points above the previous best and above the usual pass mark; their tuned model, Med-PaLM, was still judged inferior to clinicians in its long answers. Med-PaLM 2 reached 86.5 percent, and in its authors’ human evaluation physicians preferred its answers to other physicians’ on eight of nine axes. General-purpose chat models were soon answering patients’ medical questions directly, outside any regulatory pathway (see The LLM Race).

Dead End: The Benchmark

Every generation of medical AI has been validated first on a set of cases chosen to be clean, labelled and complete, and has then met patients who were none of those. The Leeds program did not travel to other hospitals, CAD improved reader studies and not screening, the eye screener rejected a fifth of the photographs taken in a Thai clinic, and the sepsis score was switched on in hundreds of hospitals before anyone outside the vendor measured it. A licensing-exam score is a benchmark of the same kind. The FDA’s list counts authorizations, not results.

📚 Sources