Most AI Scribe Rollouts Underperform. The 2026 Data Shows Exactly Why.
The largest studies to date find AI scribes save about 13 minutes of EHR time a day and do not move after-hours charting at all — while heavy users save nearly double. The difference is the rollout. Here is what the 2026 implementation evidence says about doing it right.
Almost every serious study published in the past eighteen months agrees on two things about AI medical scribes: clinicians who use them are measurably happier, and the average measured time savings is a lot smaller than the marketing promised.
Both things are true at once. The reason is not that the technology is bad. It is that most of the value sits with a minority of users, and whether a clinician ends up in that minority is decided almost entirely by how the rollout was run.
This is a guide to running one well — built from what the 2025–2026 implementation literature actually reports, including the numbers vendors do not put on their slides.
The headline result you need to internalize
On April 28, 2026, JAMA published the largest multi-site study of ambient scribes to date: "Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence–Powered Scribes: A Multisite Study" (Rotenstein LS, Holmgren AJ, Thombley R, et al.; JAMA. 2026;335(16):1408–1417).
The design is about as good as observational health IT research gets:
- 8,581 ambulatory clinicians across five academic systems — Mass General Brigham, Emory Healthcare, UCSF, Yale New Haven Health, and UC Davis
- 1,809 adopters and 6,772 non-adopters, June 2023 through August 2025
- Three commercial vendors — Ambience, Nuance DAX Copilot, and Abridge — all running on Epic
The results, normalized to eight scheduled patient hours:
- Total EHR time: 13.4 fewer minutes (95% CI, 9.1–17.7)
- Documentation time: 16.0 fewer minutes (95% CI, 13.7–18.3)
- Weekly visit volume: 0.49 additional visits (95% CI, 0.17–0.81)
- EHR time outside scheduled hours: no significant change
In relative terms that is roughly a 3% reduction in total EHR time and a 10% reduction in documentation time. Real, statistically solid, and roughly a quarter of what a typical vendor case study implies.
The line that should stop you, though, is the last row. After-hours EHR time did not move. "Pajama time" is the thing physicians actually complain about, and in the largest study we have, the scribe did not touch it on average.
The authors also found the benefit was not evenly distributed. It concentrated in primary care, advanced practice clinicians, female clinicians, and — critically — intensive users.
Reported secondary analyses of the same cohort put "power users" (those using the scribe on at least half their eligible encounters, about a third of adopters) at roughly 21 minutes of EHR time and 27 minutes of documentation time saved — close to double the headline figure.
That is the whole game. The average is mediocre because the average blends heavy users with people who touched the tool twice in March.
The same pattern shows up everywhere
Once you know to look for it, the concentration effect is in every large dataset.
The Permanente Medical Group ran what is still the largest single-organization deployment on record: 7,260 physicians, 2.5 million+ encounters over 63 weeks from October 2023 through December 2024, published in NEJM Catalyst. Total documentation time saved was about 15,791 hours — roughly 1,794 eight-hour workdays.
But the distribution is the story: frequent users accounted for 89% of all scribe activations, and they saved about 2.5× more time per note than infrequent users. Adoption was highest in mental health, emergency medicine, and primary care, and — a genuinely useful finding for anyone bracing for generational pushback — physician age showed no correlation with adoption.
Patients noticed, too. Forty-seven percent said their doctor spent less time at the computer, and 39% reported more direct conversation. Eighty-four percent of physicians said the tool improved patient interactions; 82% reported better overall work satisfaction.
Vanderbilt University Medical Center published the most honest deployment write-up I have read (Wright AP, et al., JAMIA. 2026;33(2):457–461). VUMC piloted Dragon Ambient eXperience with 54 ambulatory and 37 emergency clinicians from March to December 2024, then went enterprise-wide in a single day on January 15, 2025, to 2,400+ eligible clinicians.
Uptake was fast:
- Day 1: 233 users
- End of February: 1,000+
- March 31: 1,223 total users, 1,032 active in the prior 30 days
And yet by the end of the study window, ambient scribing was used on 20.1% of ambulatory and ED visit notes — up from 4.2%, but still one note in five, in an organization where essentially everyone had a license.
That is not a failure. It is what a good rollout looks like. Anyone promising you 80% note penetration in a quarter is selling something.
The failure mode is adoption, not accuracy
Practices spend their evaluation energy on note quality and their rollout energy on almost nothing. The evidence says that ratio is backwards. Accuracy matters — see our review of what the research says about scribe accuracy and hallucinations — but it is rarely what kills a deployment.
VUMC catalogued why clinicians never started:
- Platform mismatch — the tool was iOS-only, so Android users were locked out
- They did not know it was available
- Their workflow was already optimized — templates and dot phrases they trusted
- General reluctance toward AI
And why clinicians who started stopped:
- Editing the draft took longer than their existing workflow
- Incompatible with copy-forward documentation — a huge factor in specialties managing chronic disease
- Excessive verbosity or inaccuracy in drafts
- They preferred the quality of an in-person scribe
- Concerns about recording, from patients or the clinician
Look at that list. Exactly one item is about model quality. The rest are onboarding, communication, workflow fit, and expectation-setting — all of which are yours to control, and none of which appear in a vendor demo.
The recording concern is the one to get ahead of before go-live, not after. Consent requirements vary by state and are actively litigated; start with our guide to AI scribe consent, retention, and recording laws and the state AI disclosure rules now in force.
Training is the single largest lever — and almost nobody pulls it
The KLAS Arch Collaborative's 2026 analysis, drawn from clinician surveys at 12 Epic organizations collected since January 2026, produced the most actionable statistic in this entire field.
Among clinicians using ambient speech tools, Net EHR Experience Score by whether they felt they knew how to optimize the tool:
- Strongly agree they know how to optimize it: 89.7
- Strongly disagree: 46.7
A 43-point spread on the same technology, at the same organizations. For comparison, the gap between clinicians who felt adequately trained on the EHR overall and those who did not was about 20 points. Ambient AI is roughly twice as training-sensitive as the EHR itself.
And the gap is wide open: fewer than 25% of clinicians who have adopted AI tools say they received adequate training on how to work with AI-generated content in their workflow.
KLAS also found a ceiling worth planning around. Clinician satisfaction rises as people adopt up to about four AI tools, then flattens — those juggling five or more report no incremental gain. If you are stacking a scribe on top of coding assistance, inbox drafting, and intake automation, sequence them. Do not ship them in the same month.
Big-bang or phased?
Both work. The choice should follow your support capacity, not your ambition.
VUMC went big-bang deliberately and documented why: cohesive communication, faster value realization, and no loss of momentum between waves. The support load turned out to be modest — 34 help desk tickets across the study window, mostly access, note generation, configuration, and training questions, with no critical safety events identified. Their conclusion was that simultaneous enterprise rollout is more feasible than most organizations assume.
Mass General Brigham and The Permanente Medical Group went the other way, scaling from tight pilots — MGB from 18 pilot physicians to more than 3,000 in under two years.
The variables that actually decide it:
- Do you have an existing pilot cohort with credible champions? VUMC's ten-month pilot is what made the big-bang viable — they entered launch day with peer advocates in most departments. Skipping the pilot and going wide is not the same strategy.
- Is your license enterprise-wide or seat-limited? VUMC found universal licensing avoided excluding part-time clinicians and cut the administrative overhead of managing allocations. Seat-limited contracts force phasing whether you want it or not.
- Can your help desk absorb a spike? 34 tickets across 1,200 users is the number from an organization with a mature informatics team and a self-service training site already built.
For a small or mid-sized practice, this question mostly dissolves — you are phasing by default because you have ten clinicians. See our comparisons for small medical clinics, family medicine, and nurse practitioners.
What VUMC did that you should copy
Four decisions from the JAMIA paper are worth lifting directly.
1. Access without a productivity mandate. No requirement to see more patients. The program was framed around clinician well-being and efficiency, full stop. This is the one most systems get wrong — and PHTI's analysts specifically warned that pushing higher patient loads to capture ROI risks erasing the burnout benefit that justified the purchase.
2. Enterprise-wide licensing. Cheaper administratively than gatekeeping seats, and it removed the part-time-clinician exclusion problem entirely.
3. Clinician champions from the pilot. Peer credibility moved adoption more than any official communication. Recruit champions during the pilot with the explicit expectation that they will be visible advocates at launch.
4. Low-barrier, self-service training. A training website with optional webinars. Notably, the live webinars were sparsely attended — clinicians used the async material instead. 84.7% still rated the training experience positively. Do not build your program around mandatory sessions people will resent.
Pair that with a governance layer. The AMA's STEPS Forward "Governance for Augmented Intelligence" toolkit, developed with Manatt Health, lays out an eight-step framework covering executive accountability, working groups, policy development, vendor evaluation, implementation, oversight, and readiness. If nobody owns AI decisions at your organization, that is the gap to close before the contract, not after.
Measure the right four things — and get a baseline first
The most common analytics mistake is having no pre-deployment baseline, which makes every post-deployment number unfalsifiable.
Pull at least eight weeks of baseline data before go-live, then track:
1. Utilization rate, per clinician. Not licenses issued — percentage of eligible encounters documented with the tool. This is the number that predicts everything else. VUMC's 20.1% note penetration and TPMG's 89%-of-activations-from-frequent-users both point the same direction: segment your users into heavy, occasional, and dormant, and manage each differently.
2. Documentation time per eight scheduled patient hours. Normalizing to scheduled hours is what the JAMA team did, and it is the only way to avoid confounding by clinicians who simply worked less that month. Epic Signal and equivalents give you this.
3. After-hours EHR time, tracked separately. The JAMA study found no significant change here on average. If yours moves, you have done something genuinely better than the field. If it does not, know that early rather than discovering it in a board presentation.
4. Burnout, with a validated instrument, at a fixed interval. The single-item Mini-Z or Maslach subscales, measured at baseline and at 60–90 days. The published effects are large enough to detect in a small practice: MGB reported a 21.2% absolute reduction in burnout prevalence at 84 days; Emory a 30.7% increase in documentation-related well-being; MultiCare a 63% reduction in burnout and 64% improvement in work-life balance.
Two things not to over-index on: note quality scored by the vendor, and self-reported time savings. VUMC's clinicians estimated 6 minutes saved per encounter by survey — a plausible figure, but self-report and telemetry rarely agree, and only one of them survives a CFO's questions.
The ROI conversation, honestly
The Peterson Health Technology Institute's March 27, 2025 assessment reached a conclusion the industry has been quietly working around ever since: ambient scribes reduce burnout and cognitive load, but have not yet demonstrated a clear financial return. Health systems reported no consistent impact on patient volumes or billing accuracy. Adoption, once tools were widely available, ran 20–50%.
The JAMA data is consistent with that. Reported analyses of the cohort put the associated E/M revenue gain at roughly $167 per clinician per month — against ambient scribe pricing that commonly lands anywhere from about $100 to $600 per provider per month depending on tier and vendor. (Our 2026 pricing breakdown walks through what practices actually pay.)
So the math can work, or not, depending almost entirely on where you land in that price range and how many of your clinicians become heavy users. Which means:
- Negotiate on price aggressively. At the top of that range, the documented revenue effect does not cover the license.
- Build the business case on retention and recruitment, not on visit volume. A physician who does not leave is worth vastly more than 0.49 visits per week. Adam Landman, MGB's CIO, framed ambient documentation as "one of the most effective and impactful methods for enhancing the provider experience" — that is the honest claim, and it is a strong one.
- Do not promise volume gains to finance. The measured effect is a 1.7% visit increase. If your ROI model assumes more, it will be wrong, and the credibility cost will land on the next AI project you propose.
For a sense of what individual systems have gotten, the AHA's April 2026 market scan reported Cleveland Clinic cutting note writing and review by 14 minutes per clinician per day with Ambience, Cooper University Health Care saving 4.15 minutes per patient with Dragon Copilot, and Intermountain Health reporting a 27% reduction in time in notes per appointment among clinicians with 10 or more encounters, from April 2024 to December 2025.
A 90-day rollout plan
Days −60 to 0: baseline and setup
- Pull eight weeks of EHR telemetry: total EHR time, documentation time, after-hours time, all normalized to scheduled hours
- Administer a validated burnout instrument
- Verify device compatibility across your actual clinician population — VUMC lost Android users to an iOS-only tool
- Settle consent language, patient-facing signage, and recording retention policy (start here)
- Confirm the BAA and data handling terms (what HIPAA compliance actually requires)
- Recruit 1–2 champions per department from your pilot cohort
Days 1–30: launch
- Ship a self-service training site with 5–10 minute async modules; make live sessions optional
- State explicitly, in writing, that there is no productivity expectation attached
- Publish a single escalation path for problems
- Track utilization weekly from day one
Days 31–60: segment and intervene
- Split users into heavy (>50% of encounters), occasional, and dormant
- Interview the dormant group. You are looking for the VUMC reasons: editing burden, copy-forward conflict, verbosity, unawareness
- Run targeted optimization coaching for occasional users — this is where the 43-point KLAS satisfaction gap lives
- Tune templates and specialty-specific formatting; VUMC and others found specialty customization essential to sustained use
Days 61–90: measure and decide
- Re-pull the same four metrics and re-administer the burnout instrument
- Compare against baseline, not against the vendor's case study
- Decide on expansion, renegotiation, or exit while you still have contract leverage
If you are a small practice
Most of the above scales down, with three adjustments.
Your pilot is one clinician for two weeks. That is enough to catch a workflow conflict.
Your baseline may have to be self-reported. Have each clinician log finish time for two weeks before and two weeks after. Crude, but a real comparison beats a vendor's number.
Your leverage is that you can switch. Enterprise contracts lock in for years; a solo clinician on a monthly plan can test two tools in a quarter. Use that. Our Heidi vs. Freed comparison, the free and low-cost options guide, and individual reviews of Abridge, Suki, and Nabla are reasonable starting points, and Freed, Heidi Health, and DeepScribe all offer paths to trial without a procurement cycle.
If you are already on Epic and weighing the bundled option, our comparison of Epic's native AI charting versus standalone scribes covers that trade-off. Behavioral health practices should start with AI therapy notes for group practices instead, since psychotherapy note handling changes the privacy calculus substantially.
The takeaway
The technology works. The published effects are real and, for the clinicians who use it heavily, substantial.
But the average result across a population is set by adoption, and adoption is set by the things nobody budgets for: device compatibility checks, async training, peer champions, workflow customization, an explicit promise that nobody will be asked to see more patients, and a segmentation review at day 45 that finds the people who quietly stopped.
Vendors will sell you the tool. The 43-point satisfaction gap between clinicians who know how to optimize it and those who do not is yours to close.
Browse AI medical scribe tools in the directory, or start with what to look for when choosing one.
This article is informational only and is not legal, medical, or financial advice. Study figures are as reported in the cited publications; confirm pricing, feature availability, device requirements, and contract terms directly with the vendor before purchasing.