HIPAA & Compliance

Your AI Scribe Is Changing How You Bill. Your Payer’s Algorithm Already Noticed.

The 2026 research shows coding intensity rises after ambient AI adoption — and payers have started downcoding in response. What the Trilliant, BCBSA, and UCSF data actually says, why an AI scribe can move an E/M level at all, and the audit checklist every practice should be running.

By MedAI Directory · August 10, 2026

Nobody bought an AI scribe to change their billing. You bought one to stop charting at 9pm.

But in 2026 the evidence got hard to ignore: practices and health systems that adopt ambient documentation tend to bill more high-level visits afterward. Payers noticed first, and they responded with algorithms of their own. Cigna now automatically knocks a level off certain E/M claims from clinicians it considers outliers. Blue Cross Blue Shield published a number — $2.3 billion — and put it in a press release.

So there's a squeeze forming, and the practice is in the middle of it. Your scribe nudges your documentation one way. Your payer's software nudges the claim back the other way. And the person who signed the note is the one who answers for it.

This post covers what the 2026 data actually shows, why an ambient scribe can move a code level at all, where legitimate capture ends and upcoding begins, and the specific things a practice should be measuring right now.

The one-paragraph version

Coding intensity has risen at organizations using ambient AI, and the size of the shift is real but smaller at the individual-physician level than the headlines suggest. Some of it is almost certainly legitimate — clinicians were under-documenting work they actually did. Some of it may not be. Nobody has cleanly separated the two yet, including the researchers who published the findings. Meanwhile payers have started downcoding preemptively, more than a dozen states have passed or introduced laws restricting that practice, and federal enforcement around AI-prompted diagnosis coding got explicit in February. The practical answer is not to avoid AI scribes. It's to baseline your E/M distribution, audit against medical decision making rather than note length, and never let a suggested code become the submitted code without a human deciding.

What the data actually shows

Coding intensity moved up across every system studied

The most-cited analysis is from Trilliant Health, released March 12, 2026. Researchers looked at national all-payer claims data from 2018 to 2024 across six large hospitals and health systems that had publicly announced ambient AI scribing adoption, tracking outpatient E/M codes 99202–99215.

The direction was consistent everywhere:

  • Established-patient visits shifted toward high-intensity codes (99214–99215) by roughly 7 to 12 percentage points. One system went from 59.8% to 67.2% high-intensity; another from 40.9% to 52.7%.
  • New-patient visits shifted by 12 to 20 percentage points. The largest mover went from 60.5% to 80.0% coded at 99204–99205.
  • The pattern held across nearly every ICD-10 diagnosis chapter, not just a few complex ones.

The authors were careful, and their caution deserves repeating: they concluded the shift "likely reflects enhanced rules-based documentation" rather than fraud, and stated plainly that "the underlying causes of the changes are not clearly understood." This is a study of a correlation across a seven-year window that also contained a pandemic, a major E/M guideline rewrite, and a shift in payer mix. It is not proof that scribes cause upcoding.

The physician-level numbers are smaller than the headlines

The cleanest causal work comes from UCSF. In JAMA Network Open on January 9, 2026, Holmgren and colleagues published a difference-in-differences study of 1,565 physicians and nearly 1.2 million ambulatory encounters between January 2023 and April 2025. Of those physicians, 698 adopted an ambient scribe and 867 did not.

What adopters gained, relative to non-adopters:

  • 1.81 additional work RVUs per week — a 5.8% increase
  • 0.80 more patient encounters per week — a 2.8% increase
  • Roughly $3,044 in additional annual revenue per physician at 2025 Medicare rates
  • No increase in claim denial rates

That last point matters. If ambient documentation were producing indefensible claims at scale, you would expect denials to climb. They didn't. But the authors also flagged the obvious limitation: this is one health system, adopters volunteered rather than being randomized, and the study cannot say whether the RVU bump reflects more work delivered or more work captured.

Individual systems report similar magnitudes. Riverside Health told Healthcare Brew it had seen an 11% rise in physician RVUs since May 2024, with level-four encounters up about four percentage points and level-five up about one.

Note the gap between these figures and the Trilliant percentages. A 5.8% RVU lift per physician is a meaningfully different story from "new-patient high-intensity coding hit 80% at one health system." Both can be true; they measure different things over different windows.

Payers are attributing real dollars to it

On March 5, 2026, the Blue Cross Blue Shield Association published research with Blue Health Intelligence estimating that more aggressive, AI-enabled coding accounted for roughly $2.3 billion in additional claims spending over a two-year period — about $663 million inpatient and at least $1.67 billion outpatient.

Their sharpest finding wasn't about E/M levels at all. Analyzing tens of thousands of maternity admissions, they found a steep rise in cases coded for acute posthemorrhagic anemia — a condition that typically warrants transfusion — at facilities where many of those patients received no such treatment. As BCBSA's Dr. Razia Hashmi put it: "Something is disconnected."

That's a coding-tool finding more than a scribe finding, and it's worth keeping the two straight. Ambient scribes draft narrative notes. Autonomous coding engines assign billable codes from charts. They're different products with different failure modes, and we covered the coding side separately in AI medical coding and billing in 2026. But payers increasingly treat them as one category of risk, and your practice will be judged by the claims that come out the far end regardless of which tool produced them.

Why an AI scribe can move your code level at all

This is where most coverage gets it wrong, so it's worth being precise.

Since 2021, office visit levels 99202–99215 are selected on medical decision making or total time. Full stop. History and physical exam are no longer scored. A longer, richer, better-organized note does not by itself justify a higher level — and if your only defense in an audit is that the note got longer, you don't have a defense.

So how does an ambient scribe actually move the needle? Legitimately, through four channels:

  • Problems addressed. Clinicians routinely evaluate three or four issues in a visit and document one. A scribe that captures the whole conversation surfaces the comorbidities that were genuinely assessed.
  • Data reviewed. Ordering labs, reviewing an outside note, discussing results with another clinician, using an independent historian — all countable, all easy to forget to write down.
  • Risk. Medication management, decisions about hospitalization, social determinants that affect the treatment plan. These often live in dialogue and die there.
  • Time. Total time on the date of encounter is an alternative path to level selection, and scribes make time documentation less painful.

Every one of those is real work that was previously invisible. That's the honest case for the intensity shift, and it's a good one.

Now the failure modes, which are just as real:

  • Mentioned is not addressed. Under CPT, a problem counts as "addressed" when it's evaluated or treated at that encounter. A chronic condition listed in the assessment because the patient brought it up in passing — or because it copied forward from last visit — is not addressed. Ambient tools are very good at populating problem lists. They are not good at knowing which problems you actually managed.
  • Suggested codes carry automation bias. Abridge and Suki both surface ICD-10 and CPT suggestions alongside the draft note; Suki extends to E/M and HCC. DeepScribe offers CPT suggestions as well. Nabla has said coding is in development rather than shipped. These features are convenient and they are also the exact point where a busy clinician stops evaluating and starts accepting.
  • Hallucinated elements inflate the denominator. If the draft documents a review of systems or a data element that didn't happen, and that element contributes to your MDM calculation, you have a coding problem layered on a safety problem. We went through the accuracy literature in what the research says about AI scribe accuracy and hallucinations.

The uncomfortable version: an AI scribe makes it easier to document a level 4 and no easier to have performed one. The gap between those two is exactly what an auditor is looking for.

The other algorithm: your payer is downcoding

While practices were adopting ambient AI, payers were deploying claims-side AI. The most visible example is Cigna's Evaluation and Management Coding Accuracy policy (R49), effective October 1, 2025, which automatically reduces certain claims — 99204–99205, 99214–99215, and 99244–99245 — by one level for clinicians showing a consistent pattern of higher-level coding relative to peers. Submit the records afterward and demonstrate the level via MDM or time, and Cigna will pay the original level. Cigna has said fewer than about 1% of in-network providers would be affected at launch. Physician groups across gastroenterology, oncology, family medicine, and sleep medicine objected loudly.

Regulators and legislatures have been moving against algorithm-only downcoding since:

  • Maryland fined Cigna $80,000 on March 13, 2026 and ordered it to stop automatically downcoding E/M claims.
  • Indiana enacted HB 1271 on March 4, 2026, barring insurers from using an automated process as the sole basis to downcode a claim on medical necessity grounds without human review of the record.
  • Illinois passed SB 3114, the Transparency in Downcoding Act, requiring a person to make or review every downcoding determination under current AMA CPT guidelines.
  • Virginia enacted its own downcoding law, and the AMA reports more than a dozen states introduced bills in 2026.
  • California's medical association paused several downcoding policies pending regulatory review.

The practical consequence for a practice is that your E/M distribution is now a monitored statistic in both directions. Drift upward and you may get algorithmically trimmed. Fail to appeal and you eat the difference permanently.

The enforcement backdrop

Two federal developments frame the risk, and it's important to describe them accurately rather than dramatically.

HHS-OIG issued Medicare Advantage compliance program guidance on February 3, 2026 — the first update since 1999. It explicitly names as a risk area the practice of "querying physicians via electronic medical record platforms, including prompts generated by artificial intelligence algorithms, to add risk-adjusting diagnoses" that patients did not have or that didn't affect their care. OIG recommends pre- and post-submission audits of diagnosis data, training staff on the proper use of diagnostic prompts, and using data analytics to find providers coding at outlier rates.

In January 2026, five Kaiser Permanente affiliates agreed to pay $556 million to resolve False Claims Act allegations that they pressured physicians to add diagnoses to charts after visits — the largest MA risk-adjustment settlement to date. That case covered conduct from roughly 2009 to 2018 and had nothing to do with ambient scribes. It matters here only because it establishes what the government thinks about after-the-fact diagnosis addition, and AI-generated prompts are functionally the modern version of the same workflow.

OIG's 2026 work plan also added a review of E/M claims billed alongside minor procedures without Modifier 25, and an audit of chronic care management eligibility documentation. Neither targets AI. Both target the kind of documentation-versus-service mismatch that AI makes easier to create at volume.

What to actually do about it

None of this is an argument against ambient documentation. The time savings are real, the denial data is reassuring, and the alternative — clinicians typing notes at midnight — has its own well-documented costs. But if you're running a practice, treat coding intensity as a metric you own rather than a side effect you discover during an audit.

Before you roll out — or right now, if you already have

  • Pull your baseline E/M distribution for the trailing 12 months, by clinician, split into new and established. You cannot detect drift you never measured. If you're already live, pull the 12 months preceding go-live.
  • Benchmark against your specialty, not your gut. CMS publishes Medicare Part B utilization data by specialty; most billing systems can produce a peer comparison. Know where you sat before AI.
  • Write down who owns final code selection. If the answer is "the scribe suggests and the clinician clicks," that's a policy, and it should be a deliberate one.

At 90 days

  • Re-pull the distribution and compare. A few points of upward movement is expected and defensible. A ten-point jump in one quarter is a signal to look harder, not a win to celebrate.
  • Audit 10 charts per clinician against MDM — not against note quality. For each, ask: how many problems were genuinely addressed? What data was genuinely reviewed? What was the actual risk? Does the submitted level survive that reading without reference to how thorough the note looks?
  • Check the physical exam and data sections specifically for elements that didn't happen. This is where hallucinations concentrate.

Every quarter after that

  • Track downcodes as a distinct denial category and appeal them. Under Cigna's policy the records restore the original level; a downcode you don't appeal is a rate cut you agreed to.
  • Watch time-based billing separately. If your practice shifted toward time-based level selection after adopting a scribe, make sure the documented time reflects real time on the date of service.
  • Re-read your attestation language. The clinician who signs is responsible for the content. No scribe vendor accepts clinical or billing liability for its drafts.

Questions worth asking your vendor

Before you sign — or at your next renewal — get answers in writing:

  • Does the product suggest CPT, E/M, or ICD-10 codes? Can suggestions be turned off practice-wide?
  • On what basis is a suggested E/M level derived — medical decision making elements, time, or note characteristics?
  • Does the tool distinguish problems addressed from problems mentioned in the assessment?
  • Is the raw transcript retained, for how long, and can we retrieve it if a claim is audited two years from now?
  • What does the contract say about liability for coding derived from the tool's output?

The retention question is more important than it sounds. If a payer challenges a claim in 2028, the transcript is the closest thing you have to contemporaneous evidence of what was discussed — which also makes it discoverable. We walked through the recording and retention tradeoffs in your AI scribe is a recording device.

The bottom line

The 2026 evidence supports a modest, real increase in coding intensity following ambient AI adoption, a portion of which is legitimate capture of work clinicians were already doing and failing to document. Nobody has convincingly quantified the remainder, and the organizations doing the loudest quantifying are the ones who pay the claims.

That ambiguity is the problem. In an ambiguous environment, the practices that get hurt are the ones who can't show their work — no baseline, no periodic audit, no documented policy on who picks the code. The practices that are fine are the ones who can hand an auditor a distribution chart, a sample of MDM-based chart reviews, and a written attestation policy, and say: here's what changed, here's why, here's how we checked.

If you're still choosing a tool, what to look for in an AI scribe and our pricing breakdown cover the selection side, and you can compare options for family medicine and small clinics directly. If you're mid-rollout, why most AI scribe rollouts underperform is the companion piece to this one — add coding intensity to the metrics you track there.


This article is informational only and is not legal, compliance, coding, or medical advice. Coding rules, payer policies, and state laws change frequently and vary by jurisdiction and contract. Verify current policy details with the vendor, your payers, and a qualified coding or healthcare attorney before making billing or compliance decisions.

Tags
AI Medical ScribeE/M CodingMedical BillingComplianceAudit RiskDowncodingAmbient AI