GVC of the Day, August 3, 2026: Clinical AI Can Cause Harm—and Still Improve Physician Performance

August 3, 2026 · Clinical AI, patient safety and human–AI teaming

Clinical AI Can Cause Harm—and Still Improve Physician Performance

On NOHARM clinical consultations, AI improved physician performance—yet valuable AI recommendations were frequently left unused.

The NOHARM preprint evaluates 1,100 primary-care-to-specialist consultation tasks spanning 10 specialties. Across individual AI systems, potentially severe errors occurred in 2.9% to 24.6% of cases; a hypothetical “do nothing” comparator reached 37%. More than 80% of severe errors were omissions—important actions left out—and the four evaluated clinical retrieval-augmented generation tools outperformed the group of leading general-purpose models. Germany’s AMBOSS LiSA ranked first on severity-weighted F1 (86.15) and had the lowest severe-harm rate (2.9%), although the severe-harm differences among the four clinical RAG systems were not statistically significant.

The study also tested AI assistance in a randomized, within-participant crossover study involving 101 U.S.-licensed attending physicians. In a secondary analysis grouped by actual AI use rather than randomized assignment, the severity-weighted performance score was 52.0% with AI use versus 42.2% without it—a difference of 9.8 percentage points.

Physicians nevertheless omitted valuable recommendations that the AI had surfaced. The authors estimate that incorporating all appropriate AI recommendations would have produced combined responses outperforming both the physicians and the AI system as actually used. The next safety frontier is therefore not simply human versus AI, but how clinicians are trained to detect omissions, reconcile disagreement, challenge unsafe output and escalate uncertainty.

This is a preprint using consultation benchmarks and expert-rated management plans. It does not demonstrate improved patient outcomes, and the observed-use comparison should not be interpreted as a randomized estimate of AI’s causal effect.

Question for the audienceIf humans and AI make complementary errors, why are we still evaluating—and training—them separately?

← All GVCs