Autonomous AI coding works reliably for a defined subset of high-volume, low-complexity encounters, and it is not yet reliable enough to run unsupervised across a full case mix. That is the consistent finding across the health systems furthest along in deployment, and it is a more nuanced answer than either the "coding is solved" pitch decks or the "AI cannot code" skeptics suggest.

Why hospitals are pushing this hard right now

The driver is not primarily ambition, it is scarcity. TechTarget reported a national medical coder shortage of up to 30 percent, citing American Medical Association figures, severe enough that systems like UC Davis Health cannot fill coding roles even with competitive, fully remote compensation. When a workforce gap of that size opens, automation stops being optional and becomes the only lever left, which changes the risk calculus around deploying an imperfect technology.

What does the survey data say is actually working?

An Oliver Wyman survey on AI's impact across revenue cycle management found real, measurable results already in production, concentrated in specific functions: claims scrubbing before submission, denial prediction, prior authorization documentation assembly, and coding support for high-volume, template-driven encounter types. The pattern across the survey is consistent with what Experian Health's 2026 report found: adoption is broad, but most providers still keep human oversight in the loop rather than trusting fully autonomous output, even where volume pressure is highest.

What "autonomous coding" actually means, and where the line sits

HFMA's coverage of the transition to autonomous coding, built around lessons from Intermountain Health and OSU Physicians, frames autonomous coding not as a single on-off switch but as a confidence-scored pipeline: encounters the model codes with high confidence go straight through, encounters below a confidence threshold route to a human coder for review. The health systems furthest along report the hardest part is not the model's accuracy, it is the organizational change management: retraining coding staff for review and audit roles rather than production coding, and building trust with clinical and compliance teams that the confidence threshold is calibrated correctly and monitored continuously.

Why full automation keeps missing its own timeline

MedCity News quoted R1's Co-CEO of innovation making a point worth repeating verbatim in spirit: two to three years ago, investors were telling startups that clinical coding would be fully automated by large language models within a year, and that did not turn out to be true. Revenue cycle work is stubbornly resistant to easy fixes because:

  1. Coding requires clinical judgment on ambiguous documentation, not just pattern matching against a code list, and ambiguous documentation is common, not rare.
  2. Coding errors are expensive in a specific, asymmetric way. Under-coding loses revenue quietly. Over-coding, or coding that does not match documented clinical work, is exactly the pattern that draws a payer audit or a compliance investigation, so the cost of a false positive is much higher than the cost of a false negative.
  3. Documentation quality varies enormously by specialty and by clinician, and a model trained on clean, template-driven notes underperforms on the narrative, idiosyncratic documentation common in complex specialty care, which is precisely the segment where coding is hardest and highest-value.

A practical framework for where to automate first

  1. High-volume, low-complexity, template-driven encounters first: routine outpatient visits, standard imaging orders, common procedures with limited code variability. This is where AI performs best today.
  2. Complex, multi-diagnosis, narrative-heavy encounters last: oncology, multi-system inpatient stays, anything with significant clinical judgment about code sequencing. Keep experienced human coders here, supported by AI-assisted suggestion rather than autonomous output.
  3. Build the audit function before the automation, not after. Every system that has done this well runs continuous sampling of AI-coded claims against a human-coded gold standard, tracked as a permanent quality metric, not a pilot-phase checkbox.
  4. Reskill, do not just replace. The systems getting the best results are retraining coders into review, audit and edge-case roles, which also solves part of the retention problem driving the shortage in the first place.

The takeaway

AI in medical coding is one of the more grounded, less hyped corners of healthcare AI in 2026, precisely because the coder shortage forced real deployment rather than pilot theater. The evidence says it is a genuinely strong assistant for the easy half of the case mix and a genuine risk if pushed to run the hard half unsupervised. Systems that respect that line will get the labor relief they need. Systems that do not will meet the difference in their next payer audit.