Ambient AI can reduce documentation time, but occasional errors in diagnoses, medications, and missing clinical information leave physicians to decide how much verification is enough.
An AI scribe documented that a patient had diabetes when the patient did not. Another produced an aspirin recommendation that was not part of the physician’s treatment plan. A third misstated a metformin instruction.
Those were among the errors identified in a UC Davis Health evaluation of AI-generated clinical notes.
Physicians reviewed 356 notes and found that 94.7% had no significant errors. But 5.3% contained an error physicians judged could pose a risk of serious harm if left uncorrected. Researchers also identified accidental omissions, hallucinations, and information inserted into notes that did not belong there.
That 5.3% figure did not establish actual patient harm. Instead, it measured the potential consequence if the error escaped correction.
That combination points to an emerging challenge in ambient AI: when most notes look accurate, the occasional consequential mistake may become harder to spot.
The technology also offers measurable benefits.
A 2026 JAMA study followed 8,581 clinicians at five U.S. academic medical centers, including 1,809 who adopted AI scribes. Adoption was associated with about 16 fewer minutes of documentation time and 13.4 fewer minutes of total EHR time per eight scheduled patient hours. Researchers found no statistically significant overall reduction in after-hours EHR use.
For physicians already burdened by documentation, even modest savings can matter.
But the AI does not own the medical record.
Chief medical information officers interviewed by Becker’s generally described clinicians as the final checkpoint for what enters the medical record. At the same time, they assigned health systems responsibility for safe deployment and vendors responsibility for their technology's reliability.
That changes the physician’s job without eliminating it.
Instead of writing every sentence, clinicians must determine whether the AI accurately captured diagnoses, medications, doses, assessments, treatment plans, and clinically important context.
In the UC Davis study, edited data were available for 960 AI-generated notes. Fourteen point nine percent were left entirely unedited. That does not establish that physicians failed to review those notes.
It does, however, raise an important operational question: how much review is necessary when AI output is usually correct but occasional errors may carry serious consequences?
Not all errors deserve equal weight.
A formatting problem is different from an incorrect drug dose. An awkward phrase differs from a false diagnosis or an omitted treatment plan.
Health systems may therefore need to look beyond broad measures of note quality and track what matters clinically: incorrect diagnoses, medication discrepancies, omitted findings, fabricated recommendations, and errors discovered after a note has been signed.
Researchers at ambient-AI vendor Suki have also questioned whether traditional documentation-quality measures adequately detect generative-AI errors and omissions. Their analysis deserves independent validation because Suki sells the technology, but it raises a useful question: Can a note look polished and still be clinically wrong?
Patient concerns reinforce the issue. Healthwatch England recently reported complaints about inaccurate AI-generated records. It found that 69% of surveyed adults said they would feel more comfortable with AI scribes if healthcare professionals clearly committed to checking the output for accuracy.
For U.S. physicians, however, the central issue is no longer simply whether ambient AI works.
It is how the technology should be governed.
Health systems routinely measure how much documentation time AI saves. They may also need to measure how much physician time goes into reviewing the draft, how often clinically significant corrections are required, what kinds of errors occur, and how often mistakes reach signed records.
The question is becoming increasingly difficult to avoid:
If most AI-generated notes are accurate but a small proportion contain potentially serious errors, how should physicians verify the medical record without giving back the time AI was supposed to save?
