
August 13, 2026 · Clinical AI deployment, workflow surveillance and patient safety
The Benchmark Ends Where Clinical AI Deployment Begins
Stanford Medicine’s experience with ChatEHR exposes a fundamental limitation of conventional AI evaluation: a benchmark tests anticipated tasks, while open-ended clinical AI enables uses its developers may never have imagined.
The deployment included seven automations and 1,075 trained users, who engaged in 23,000 sessions during the first three months. Clinical summarisation was the most common task.
For summaries generated after deployment, monitoring of a 10% sample estimated an average of 0.73 hallucinations and 1.60 inaccuracies per generation. These measurements matter—but model output is only one part of clinical safety.
My takeaway: once AI enters the medical record, surveillance must examine the complete clinical workflow:
- What clinicians ask the AI to do;
- What patient information is retrieved, omitted, or misunderstood;
- Whether generated claims are supported by the record;
- What clinicians accept, edit, reject, or act upon;
- How decisions, referrals, and workload change;
- What benefits, harms, and near misses occur;
- Whether performance changes across populations, specialties, or model updates.
Generative clinical AI requires surveillance capable of discovering both unanticipated uses and unanticipated harms. A fixed benchmark cannot adequately evaluate workflows that emerge only after deployment.
This is an implementation Comment supported by descriptive deployment data—not a controlled evaluation demonstrating improved clinical outcomes.
GVCs are Grains of Vital Cognizance, by Prof. Georgi V. Chaltikyan, MD, PhD.
