AI generated illustration of a hospital workstation at night with a clinician typing at a keyboard
← Back
AI in Medicine · Long Read

Medical AI Enters the Real World: The Results Are More Complicated Than Expected

September 7, 2026 · 12 min read

Quick takeaways

  • In a Kenyan primary-care trial, treatment failure was 2.2% with LLM support and 2.0% without it.
  • AI prioritization did not shorten the 53-day median wait for CT in the LungIMPACT trial.
  • A mammography study cut radiologist workload by 63.6%, though recalls increased.
  • The LiON study found 15 malignancies that had been overlooked, but it did not include a simultaneous control group.

Medical AI has spent years looking good on test sets. A model reads stored scans or answers medical questions, and researchers compare its score with a doctor’s. That setup can measure accuracy. It cannot show what happens when the same system enters a busy clinic.

Hospitals add problems that are missing from a benchmark. Patient histories can be incomplete. A warning can sit in the electronic record when nobody owns the next step, or when the next scan cannot be scheduled.

Several large studies published in 2026 tested medical AI inside real clinical workflows. I found the mixed results more useful than another accuracy benchmark. Some systems reduced workload or caught overlooked findings. Others completed their assigned task without changing the patient’s care.

The benchmark era was never enough

A retrospective study uses data that already exist. A prospective trial follows the system during actual use. A randomized trial compares care with and without the intervention. I give the later designs more weight when the claim is about patient care.

The LungIMPACT trial shows the problem. AI moved suspicious chest X-rays up a worklist, yet the median time to CT stayed at 53 days. The faster signal did not clear the delay later in the pathway.

Sensitivity, specificity and area under the curve measure prediction. For a clinical product, I also want to know what happened after the prediction: whether anyone acted on it and whether the patient’s care changed.

A primary-care LLM looked safe but did not improve the main outcome

Researchers at 16 primary-care facilities in Kenya randomized clinical officers to use an electronic medical record with or without help from a large language model. The trial included 9,691 patients and tracked treatment failures over the next 14 days (Agweyu et al.).

Treatment failure occurred in 2.2% of the AI-assisted group and 2.0% of the control group. The difference was not statistically significant. Investigators also found no serious safety signal related to the intervention.

The LLM looked safe in this trial, but it did not improve the main outcome. I care about that result more than an exam score because it came from actual patient care.

The trial tested one system and one workflow, using a relatively uncommon short-term outcome. A smaller benefit may have been hard to detect. If the product is supposed to improve treatment, though, treatment failure is a stronger test than the quality of its generated answers.

Ninety-three thousand chest X-rays and no faster cancer pathway

The LungIMPACT trial tested AI prioritization on 93,326 primary-care chest X-rays in the United Kingdom. Images marked as suspicious could move up the radiology worklist. Researchers measured whether that changed the time to CT and lung-cancer diagnosis (Woznitza et al.).

Median time to CT was 53 days in both groups. Lung cancer was diagnosed after a median of 44 days with prioritization and 46 days without it, a difference that was not significant. Neither treatment timing nor stage at diagnosis improved.

The algorithm changed the worklist. It did not shorten the cancer pathway. CT capacity and the steps after the X-ray still controlled how fast the patient moved.

The AI and radiology reports also disagreed often. Expert review found actionable findings in some of those cases. The system produced useful information, although the trial did not show a faster cancer pathway.

Mammography showed what a narrower job can accomplish

A breast-screening trial in Spain prospectively evaluated 31,301 women using standard double reading and a partially autonomous AI strategy. In the AI pathway, examinations classified as low risk were treated as normal without human reading. Radiologists reviewed the remaining cases with AI support (Elías-Cabot et al.).

Radiologist workload fell by 63.6%. The cancer-detection rate rose from 6.3 to 7.3 cancers per 1,000 screenings. I find this result more convincing because the model had one defined job inside the screening process.

Recalls increased overall, especially in digital mammography. That can send more people without cancer through additional imaging or biopsies. The results also differed between ordinary digital mammography and three-dimensional tomosynthesis.

I would track the cancer-detection rate and the recall rate together. Reporting only the extra cancers would leave out what happened to patients who were called back and did not have cancer.

The liver study points toward AI as a safety net

Researchers used an AI system called LiON to analyze contrast-enhanced CT scans for liver malignancies. After multicenter retrospective validation, they deployed it as an additional reader in routine practice across 10,333 patients (LiON investigators).

The system found 51 previously overlooked lesions, including 15 malignancies. Radiologists amended 37 reports. Twenty-two cases went to multidisciplinary teams, and clinical management changed for a subset of patients.

I like the additional-reader setup. The radiologist reads the scan first, and LiON checks for a missed liver lesion. The 15 malignancies show what the second read added.

The study was single-arm, so it had no simultaneous control group. It also did not establish a survival benefit. Comparative trials across more health systems would give a clearer measure of LiON’s effect.

Did the AI result reach the patient? Across four clinical studies, two systems changed workflow or detection while two did not change the measured patient pathway. Did the AI result reach the patient? Kenya LLM · 9,691 patients 2.2% AI vs 2.0% control · NO CHANGE LungIMPACT · 93,326 X-rays CT 53 vs 53 days · diagnosis 44 vs 46 days · NO CHANGE Spain mammography · 31,301 women workload −63.6% · detection 6.3 to 7.3 per 1,000 · CHANGED LiON · 10,333 patients 51 lesions · 15 malignancies · CHANGED
The result reached the workflow in two of the four studies. (Agweyu et al.; Woznitza et al.; Elías-Cabot et al.; Kuo et al.)

Workflow is part of the medical intervention

The software and the workflow have to be judged together. An alert needs someone responsible for reading it and a clear next step. Otherwise, the model can be correct without changing the patient’s care.

Clinicians may ignore a tool that creates too many alerts. They may distrust a score they cannot explain or use the system differently during a busy shift. A model-performance table does not capture those reactions.

I would not use workload reduction as the only measure of success. A program should also report patient outcomes and false alarms.

What should qualify as proof?

The evidence should match the claim. A diagnostic study can support a claim about cancer detection. Claims about faster treatment or better health require treatment-time data or patient outcomes.

Some tools would require years or enormous sample sizes to show a mortality effect. Documentation time and missed findings can still be useful outcomes. The product’s evidence should support the outcome being advertised.

Results can also change by hospital. Equipment and patient populations vary. I would want a health system to test the model locally and keep checking its performance after deployment.

My Thoughts

I care more about these clinical trials than another benchmark. They show where the model enters care and whether anything changes after it gives an answer.

The primary-care LLM did not reduce treatment failure. LungIMPACT did not shorten the time to CT. Those results should affect how hospitals judge similar products.

My standard is what changed for the patient after the algorithm gave its answer. In the liver study, the second read found 15 overlooked malignancies. In the mammography study, workload fell and cancer detection rose, while recalls also increased. Those are results a hospital can measure again after deployment.

Post on XShare on LinkedIn

Sources & References

  1. Agweyu A et al. “Generative AI-Enabled Clinical Decision Support System in Primary Care: A Pragmatic, Cluster-Randomized Trial.” Nature Medicine. 2026.
  2. Woznitza N et al. “AI-Based Chest X-Ray Prioritization in the Lung Cancer Diagnostic Pathway: The LungIMPACT Randomized Controlled Trial.” Nature Medicine. 2026.
  3. Elías-Cabot E et al. “AI-Based Triage and Decision Support in Mammography and Digital Tomosynthesis for Breast Cancer Screening.” Nature Medicine. 2026.
  4. “Large-Scale AI-Guided Liver Malignancy Diagnosis: Multicenter Study and a Single-Arm Trial.” Nature Medicine. 2026.

Who’s writing this?
A Loyola High School student writing about medical AI, nutrition, and the parts of health research that are easy to oversimplify.

Read Bio →