Your Patient Inbox Grew 153%. AI Drafts Are Not Fixing It — The Routing Is.
JAMA put a number on the inbox problem in June 2026: patient portal messages up 153% since 2020, office visits up 17%. The industry answer is AI draft replies — and four studies now show they save physicians almost no time, help nurses far more than doctors, and get worse as messages get longer. Here is the workflow that actually moves turnaround time, plus the e-visit billing trap and the disclosure rule that review-before-send quietly solves.
Ambient scribes got the attention because the visit is where clinicians felt the pain first. But the visit has a natural ceiling — you can only see so many patients. The inbox does not. It grew while nobody was staffing for it, and in June 2026 JAMA finally put a number on how much.
Patient-written portal messages rose 153% between 2020 and 2025, from 0.99 to 2.50 messages per patient per year. Office visits over the same period rose 17%. Telephone encounters fell about 6%.
Read those three numbers together and the conclusion is uncomfortable: messaging did not replace anything. It was added on top.
The industry's answer has been generative AI draft replies, now one of the most widely deployed AI features in healthcare after ambient documentation. Four separate studies have now measured what those drafts actually do. The honest summary is that the drafting is not the hard part — and the practices getting value out of it are the ones that redesigned the routing around it.
This is a guide to that redesign.
What the JAMA data actually says
Long JJ, McAdams-DeMarco MA, Schwartz MD, Chodosh J, Oermann EK, Segev DL, Mankowski MA. "Trends in Patient Portal Messages, Office Visits, and Telephone Encounters." JAMA. 2026;336(3):252–254. doi:10.1001/jama.2026.8690 (published online June 22, 2026).
The study analyzed Epic Cosmos data — over 8 billion encounters across roughly 2,000 hospitals and tens of thousands of clinicians — covering 2020 through 2025.
- Patient-authored portal messages: 0.99 → 2.50 per patient per year (+153%)
- Clinician- and staff-authored messages: 4.59 → 5.70 per patient per year (+24%)
- Office visits: +17%
- Telephone encounters: roughly −6%
- Messaging was most common among women, patients aged 40–64, and patients in neighborhoods with lower social vulnerability
The second bullet is the one that gets skipped. For every message a patient sends, your practice sends more than two. The inbox is not a one-way intake queue; it is a conversation, and conversations do not close in a single reply.
An accompanying editorial made the practical point plainly: health systems need to allocate scheduled time for inbox work. Otherwise the work lands after hours, or it does not get done.
Four studies, one consistent pattern
Here is where the evidence gets interesting, because it does not say what vendor marketing says.
1. Drafts did not reduce reply time. Tai-Seale M, et al. JAMA Network Open. 2024;7(4):e246565. A randomized waiting-list quality improvement study of primary care physicians at an academic health system. Physicians with AI drafts spent 21.8% more time reading messages and showed no significant reduction in time spent replying. The reported benefit was cognitive: starting from an empathetic draft rather than a blank box.
2. Almost nobody used them — except nurses. English E, Laughlin J, Sippel J, DeCamp M, Lin CT. "Utility of Artificial Intelligence–Generative Draft Replies to Patient Messages." JAMA Network Open. 2024;7(10):e2438573. UCHealth deployed GPT-4 drafting inside Epic across nine clinics and 166 users — 12 nurses, 14 medical assistants, 93 clinicians.
- 21,323 drafts generated. 2,596 used. That is 12%.
- 92% of nurses agreed it improved efficiency, empathy, and tone
- Net Promoter Score: nurses +58, medical assistants −29, clinicians −43
- Only 12% of clinicians agreed that significant edits are rarely needed
Same tool, same health system, same week. A 101-point NPS spread between the nursing staff and the physicians.
3. Better prompting moved the needle a little. Mandal S, Nov O, Wiesenfeld BM, et al. "Utilization of generative AI-drafted responses for managing patient-provider communication." npj Digital Medicine. 2025;8:591. A retrospective audit-log study of 75 health care professionals and more than 55,000 messages at a large New York City health system, October 2023 – August 2024.
- Overall draft utilization: 19.4% (rising from about 12% to 20% as prompts were refined)
- Turnaround time improved 6.76% — 331 seconds with a draft versus 355 seconds without
- Physicians preferred short, information-dense drafts; support staff preferred warmer ones
- Drafts were generated for every message, including the ~80% that never received a reply — pure added review burden
Twenty-four seconds per message is real. It is also not the number anyone bought the tool for.
4. Long drafts cost more than they save. Seegmiller P, Gatto J, Greer SE, Isingizwe GB, Ray R, Burdick TE, Preum SM. "How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting." Presented at the 2026 Annual Meeting of the Association for Computational Linguistics (July 2026). Dartmouth researchers compared clinician replies against drafts from Claude, Gemini, ChatGPT, Llama, Aloe, and Qwen across 146,000 conversations involving 10,105 patients at a large rural health system.
- Drafts needed to require less than 30% editing to deliver substantial benefit
- At roughly 75% editing, the physician is spending more time and energy than writing from scratch
- Failure modes were consistent: overly long answers, missing follow-up questions, irrelevant or inaccurate clinical detail
- Adapting models to individual clinician communication style improved accuracy 33% and cut editing 26%
- Short messages showed roughly 25% time savings; long drafts often required extensive revision
A systematic review — Hu D, Guo Y, Zhou Y, et al., npj Health Systems. 2025;2:27 — pulled together 23 studies published from 2023 to 2025, all US-based, and reached a compatible conclusion: quality and empathy comparable to human-drafted replies, inconsistent performance, and unresolved safety and transparency concerns.
The pattern across all four: AI drafting helps on short, routine, low-stakes messages, and it helps the staff layer far more than it helps physicians. It does not help on the long, ambiguous, clinically loaded messages that actually consume a physician's evening. Which means the question is not should we turn drafts on — it is which messages ever reach a clinician in the first place.
The safety finding you cannot design around with better prompts
Biro JM, Handley JL, McCurry JM, et al. "Opportunities and risks of artificial intelligence in patient portal messaging in primary care." npj Digital Medicine. 2025;8:222. doi:10.1038/s41746-025-01586-2.
Twenty practicing primary care physicians reviewed 18 patient portal messages with AI-generated drafts in a simulated EHR. Four drafts contained seeded errors — objective inaccuracies or potentially harmful omissions.
- Each error was missed by 13 to 15 of the 20 physicians
- Only one physician caught all four
- Participants missed an average of 2.67 of 4 errors
- 35% to 45% of the erroneous drafts were submitted completely unedited
Now the same participants' self-assessments:
- 80% agreed the drafts reduced cognitive workload
- 90% agreed they trusted the tool's performance
- 75% agreed the drafts were safe to use
- 70% agreed the drafts were accurate
That gap is the finding. Confidence in these tools rises faster than accuracy does, and the reviewers who were most comfortable were not most correct. This is textbook automation complacency, and it is the same mechanism behind the hallucination risk documented in what the research says about AI scribe accuracy — except that a bad scribe note gets caught at signing, while a bad portal reply is already in the patient's hands.
A follow-on paper — Sakr J. "Guardrails for GenAI drafted replies in patient portal messaging." npj Digital Medicine. 2026;9:283 — argues for exactly the response the evidence implies: scoped use, risk tiering, accountable human authorship, auditability, and patient-facing transparency. Not better models. Better process.
The workflow: seven decisions, in order
This is the part vendors do not sell you, and it is the part that determines whether any of this works.
1. Split administrative from clinical at the door
Most of your inbox is not medical decision-making. It is refill requests, scheduling, forms, billing questions, and results follow-ups. Route those away from the clinician queue before an AI ever drafts anything.
Conveniently, California's AI disclosure law already draws this line for you in statute (more below): administrative matters — scheduling, reminders, billing, clerical business — are treated as a different category from patient clinical information. Use the same boundary in your routing rules. Tools built for AI patient communication, patient intake, and scheduling are designed for this layer specifically.
2. Turn drafting on for the staff pool first
The UCHealth NPS data is unambiguous: nurses +58, clinicians −43. If you enable drafting for everyone at once, you will conclude the tool failed, because the loudest voices in the room will be the ones it helps least.
Enable it for nursing and support staff. Measure for a month. Then decide about clinicians.
3. Cap it by message length and complexity
The Dartmouth 30% rule is the most actionable number in this literature. If a draft needs more than roughly a third rewritten, it is costing time. Long, multi-question, ambiguous messages predictably generate long, wrong drafts.
Configure the tool — or instruct the team — to skip drafting on messages above a length threshold or flagged as multi-issue. Human-first is the correct default for those.
4. Read the patient's message before the draft
Small habit, real effect. The Biro errors were missed because reviewers anchored on a fluent, confident draft. Reading the original message first gives you an independent expectation to check the draft against, rather than a plausible answer to nod along with.
5. Nothing clinical auto-sends. Ever.
One named human is the author of record for every clinical reply. This is not only a safety position — it is the legal safe harbor and the patient expectation, as the next section shows.
6. Put inbox time on the schedule
The JAMA editorial's point stands whether or not you adopt AI. Unscheduled work becomes after-hours work. If drafting saves 24 seconds a message and you handle 60 messages a day, that is 24 minutes — worth having, and worth protecting on a calendar rather than donating back to volume.
7. Measure turnaround and after-hours minutes, not draft acceptance
Vendors report draft acceptance rate because it is the number that looks best. It measured 12% at UCHealth and 19.4% at NYU Langone, and neither tells you whether your clinicians got their evenings back. Track median turnaround time, after-hours EHR minutes, and unanswered-message backlog. Those are the outcomes.
The billing math, and the trap inside it
Portal messages can be billed as online digital E/M services — CPT 99421, 99422, and 99423 — when the message is patient-initiated, the patient is established, and the reply requires the clinician's medical decision-making, measured as cumulative clinician time across a seven-day window.
2026 Medicare national non-facility amounts:
- 99421 (5–10 minutes) — 0.47 total RVUs, about $15.70
- 99422 (11–20 minutes) — 0.92 total RVUs, about $30.73
- 99423 (21+ minutes) — 1.46 total RVUs, about $48.77
Standard exclusions apply: not billable if it relates to an E/M service in the prior seven days, not billable if it leads to a visit within the following seven days, and not billable for non-evaluative traffic like results delivery, scheduling, or billing questions. Verbal consent for communication-based technology services is required annually, and patient cost-sharing applies.
Then there is the reality check. Dunlay SM, et al. "Implementation of Billing for Patient Portal Messages as E-visits in a Large Integrated Health System." Annals of Internal Medicine, published December 31, 2024 — Mayo Clinic implemented e-visit billing on August 18, 2023 and studied the six months that followed:
- 5,183 medical advice messages were billed. That is 0.3%.
- Patient-initiated medical advice message threads fell 8.8% versus the same period a year earlier
- No difference in 7-day emergency service use between patients who sent a message and those who did not after seeing the billing disclaimer
Billing did modestly reduce volume, and it did not appear to push patients toward the ED. But 0.3% is not a revenue strategy.
Here is the trap. These are time-based codes, and 5 minutes of cumulative clinician time is the floor. As covered in whether Medicare pays for AI, the CY 2026 fee schedule's 2.5% efficiency adjustment deliberately excluded time-based services — but on this particular code family, efficiency reduces your revenue directly and without any help from CMS. A message that took you seven minutes and paid $15.70 takes four minutes with a good draft and pays nothing.
Two consequences:
- Do not build the ROI case for inbox AI on e-visit billing. The case is turnaround time, after-hours burden, and staff capacity. Those are the returns that survive.
- If you do bill, the time you record must be actual clinician time. Review time counts; the two seconds the model spent generating does not. Inflating a four-minute review into a five-minute claim to preserve a $15.70 line is precisely the kind of pattern payer analytics are now built to find — the same dynamic described in AI scribes and coding intensity.
What you have to tell patients
California AB 3030, codified at Health & Safety Code §1339.75 and effective January 1, 2025, is the most specific law in the country on this exact use case. It applies to any health facility, clinic, physician's office, or group practice office that uses generative AI to produce written or verbal patient communications pertaining to patient clinical information.
If it applies, the communication must carry a disclaimer that it was generated by AI:
- Written: prominently at the beginning of each communication
- Chat: prominently displayed throughout the interaction
- Audio: verbally at the start and the end
- Video: prominently displayed throughout
It must also give clear instructions for how to reach a human. Administrative matters — scheduling, billing, clerical business — are excluded. Enforcement runs through the Medical Board of California and the Osteopathic Medical Board for physicians, and existing regulatory mechanisms for facilities.
And then the sentence that should shape your entire workflow: if the AI-generated communication is read and reviewed by a human licensed or certified health care provider, the disclaimer requirement does not apply.
Review-before-send is not just the safety answer and the billing answer. In California, it is also the compliance answer. If you are practicing elsewhere, check your own state — the disclosure landscape is moving fast, and state AI disclosure laws in healthcare tracks where it stood going into 2026.
Patients, for their part, have told researchers roughly the same thing. Owens K, Jayaram A, Chowdhury A, et al. "Patient Perspectives on AI-Drafted Electronic Portal Messages." JAMA Network Open. 2026;9(7):e2622463, published July 7, 2026, interviewed 40 patients at a large academic health system:
- Comfort with AI drafting was high but conditional on clinician review and approval
- Broad support for disclosure, described as trust-building, though preferred timing and format varied
- Tone should scale with stakes — brevity for routine requests, longer and more supportive replies for ambiguous test results
- Patients viewed portal messaging as transactional and efficiency-focused: "I'm looking for quick information… I don't have time for this long thing."
That last point contradicts the default behavior of every general-purpose model, which pads. It is also consistent with the Dartmouth finding that overly long drafts are where the editing burden lives. Shorter is not a compromise here. Shorter is what patients asked for.
One more baseline: whatever tool drafts your replies is handling PHI, which makes it a business associate and makes a signed BAA non-optional. What HIPAA compliance actually means for AI tools covers what to verify, and the consent and recording rules piece covers the adjacent disclosure obligations.
Where the tools sit today
If you are on Epic, Oracle Health, or another major EHR, draft replies are most likely already available as a module — and the practical calculus there is the same one described in Epic's AI scribe versus standalone tools: the native option wins on integration, not on quality.
For the layer before the clinician queue — the routing, deflection, and administrative traffic that should never reach a physician — the directory tracks several platforms:
- Klara consolidates text, web chat, voicemail transcription, and secure messaging into one queue, with Klara Assist handling routine outreach. Custom pricing; reports suggest plans starting around $300/month.
- Hyro is built for health system call centers and patient access across phone, SMS, web chat, and portals. Enterprise contracts, typically $10,000–$50,000+/month.
- Notable automates intake, prior authorization, scheduling, and patient communication as one workflow layer. Enterprise contracts.
- Phreesia owns the intake and check-in side, handling roughly 1 in 7 US patient visits, typically from around $250–$300/month.
Be realistic about fit. Three of those four are enterprise-priced, which is a poor match for a solo or small practice — and the honest recommendation for a small practice is to exhaust what your EHR already includes before adding a platform. Side-by-side views: patient communication tools for small medical clinics (five tools), for family medicine physicians, and patient intake for small clinics. If phone volume is the bigger problem than the portal, start with AI for the front desk.
Adopt this in four weeks
- Week 1 — Measure. Pull message volume, median turnaround, after-hours EHR minutes, and the administrative-versus-clinical split. You cannot evaluate a change you never baselined. The same discipline that makes scribe rollouts succeed applies here.
- Week 2 — Route. Move administrative traffic out of the clinician queue. This is the single largest win available, and it requires no AI at all.
- Week 3 — Enable drafting for staff only. Nurses and MAs. Cap drafting on long or multi-issue messages. Instruct everyone to read the patient's message before the draft.
- Week 4 — Set the rules in writing. Nothing clinical auto-sends. One named author per reply. Scheduled inbox blocks. Annual verbal consent captured if you intend to bill e-visits. Disclosure language on file if any AI-generated clinical communication ever leaves without provider review.
Then re-measure against the week-1 baseline. If turnaround and after-hours minutes have not moved, the problem is your routing, not your model.
The honest summary
The strongest evidence we have says AI draft replies save a physician somewhere between nothing and 24 seconds per message, help nurses considerably more than physicians, and get worse as messages get longer. It also says that when a draft contains an error, most physicians miss it — while reporting high trust in the tool.
None of that makes drafting a bad idea. It makes drafting the last thing you should configure, not the first.
Your inbox grew 153% in five years, and the thing that grew fastest was the share of it that never needed a clinician. Route that away, put drafts where the data says they work, keep a named human on every clinical reply, and stop counting draft acceptance as a result.
The model is not the intervention. The workflow is.
This article is informational only and is not legal, billing, or medical advice. Coding rules, state AI disclosure law, Medicare payment amounts, and vendor capabilities change frequently — verify current requirements with CMS, your MAC, your state medical board, your compliance counsel, and the vendor before making billing, disclosure, or purchasing decisions. Payment amounts shown are national averages and vary by geographic locality and payer.