Ambient documentation is the first clinical AI category to cross from pilot to enterprise standard. Large systems have moved from dozens of licensed clinicians to thousands within a single budget cycle, which is a speed of adoption no prior clinical software has achieved. The interesting question is no longer whether it works. It is what "works" has actually been measured to mean.

What the published evaluations consistently show

Across the health-system evaluations published so far, three findings repeat with enough consistency to be treated as robust.

Clinician-reported burden falls. Measures of documentation burden and emotional exhaustion improve, often substantially, and the effect is largest in the specialties with the heaviest narrative load: primary care, behavioural health, and outpatient subspecialty clinics with long histories.

Time in the note falls, particularly after hours. The "pyjama time" metric, minutes spent in the EHR outside scheduled hours, is where the effect is cleanest, because it is measured by the EHR itself rather than by asking a tired clinician how they feel.

Attrition of the intervention is high in a predictable pattern. A significant fraction of clinicians stop using the tool within the first months, and the ones who stop tend to be procedural specialists and high-volume clinicians whose notes are already templated. The tool helps most where the note is a story and least where the note is a form.

What the evidence does not yet establish

Here is where the market is running ahead of the data.

Throughput. Very few evaluations demonstrate additional patients seen per session. Most systems that bought on a productivity thesis got a wellness result instead. That may still be worth the money, given the cost of replacing a departing physician, but it is a different business case and it should be underwritten as one.

Note quality and downstream harm. Ambient notes are longer and more complete. Longer is not automatically better. A note that faithfully transcribes an entire conversation into the chart transfers the summarisation burden to the next clinician who reads it, and there is early signal that specialists are spending more time reading primary-care notes than before. Nobody has properly measured the net effect on the system.

Coding and revenue integrity. Ambient output feeds coding. If documentation richness increases without a change in the underlying clinical work, the coding profile shifts, and that is precisely the pattern payers audit. Any system deploying at scale without a coding-integrity review is accumulating a compliance liability quietly.

Hallucination in the clinical record. The failure mode is not gibberish, it is plausible fabrication: a negated symptom rendered as present, a family history attributed to the patient, a plan detail that was discussed and then reversed. Attestation shifts legal responsibility onto the clinician, which is appropriate, but attestation quality falls as trust in the tool rises. The systems with mature governance sample notes for audit continuously rather than only during the pilot.

How to run an evaluation that tells you something

If you are deploying this year, four design choices separate a real evaluation from a testimonial.

Set the primary endpoint before go live, and make it one thing. Documentation time after hours is the most defensible. Choose it or choose retention, but do not claim both retrospectively.

Instrument a matched comparison group. Enthusiastic early adopters improve on every metric regardless of the intervention. Without a comparison arm you are measuring enthusiasm.

Audit note accuracy on a rolling sample, blinded, with a clinician reviewer and a fixed error taxonomy. Track it as a quality metric forever, not as a pilot deliverable.

Track discontinuation as your headline adoption number rather than licences issued. Licences are a procurement statistic. Sustained weekly use at month six is the only number that predicts value.

Why this category still matters

I am deliberately hard on the evidence here, and I still think ambient documentation is the most important clinical AI deployment of the decade so far. Not because of the minutes saved, but because of what it establishes structurally.

It is the first time health systems have put a generative model into the clinical workflow of thousands of clinicians, built the consent and governance scaffolding around it, negotiated the data terms, and survived the audit questions. Every subsequent clinical AI product will travel down the road this category paved: same procurement committee, same security review, same attestation habit, same integration surface.

The scribe is not the destination. It is the on-ramp, and the systems learning to evaluate it properly are the ones that will be able to evaluate everything that follows.