FDA's Generative AI Device Paper: What It Means for the AI Tools Your Practice Uses (Comments Due Oct. 19)
On August 18, 2026, FDA published its first detailed sketch of how it might regulate generative AI medical devices: a two-axis risk map, clinician-style competency testing, and more postmarket monitoring. It is not guidance and binds no one yet. Here is what it says, where scribes, CDS, patient chatbots and imaging tools land, what the first 'patient-facing LLM' clearance actually covered, and which of the 26 questions clinicians should answer before the October 19 deadline.
When the FDA rewrote its clinical decision support guidance in January, the biggest gap was generative AI. The guidance said almost nothing about chatbots, open-ended language models, or patient-facing AI.
On August 18, 2026, FDA started to fill that gap. The Digital Health Center of Excellence in FDA's device center (CDRH) published "Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback." It lays out how FDA might sort generative AI tools by risk, how it might test them before they reach patients, and how it might watch them after launch.
Comments are open under docket FDA-2026-N-7874 until October 19, 2026.
This post covers what the paper says, what it means for the AI tools small practices actually use, and why a practicing clinician might want to comment.
What this document is — and is not
Start with FDA's own caveat, because a lot of the coverage has skipped past it. The paper says it:
"is intended for discussion purposes only and does not represent draft or final guidance."
FDA also says the paper is not meant to propose or carry out policy changes, not meant to set evidence expectations for future submissions, and not meant to decide whether FDA has the legal authority for any of it.
So nothing in it binds anyone today. What it does do is show where CDRH's thinking is heading. The paper ends with 26 numbered discussion questions, and FDA says the answers "may inform the development of further policy in this area." It gives no timeline for that policy.
The paper builds on two earlier FDA advisory committee meetings. The Digital Health Advisory Committee discussed lifecycle issues for generative AI devices on November 20–21, 2024. It met again on November 6, 2025 about generative-AI mental health devices, working through a hypothetical prescription LLM therapy chatbot for major depressive disorder.
Idea 1: A two-axis map of risk
The core of the paper is a simple chart. FDA calls it "a possible organizing heuristic," not a rule.
- Horizontal axis: what the software does. It runs from non-directive information, to information that directs an action, to action with clinician supervision, to fully autonomous action.
- Vertical axis: what happens if the output is wrong. It runs from limited harm to severe harm.
Risk rises toward the upper-right corner. FDA's examples show both axes at work:
- An incorrect suggestion about an over-the-counter remedy for a minor symptom is very different from an incorrect suggestion about adjusting an insulin dose.
- Telling a parent whether to take a child with a sore throat to the pediatrician is lower risk than telling a patient whether to go to the emergency department for chest pain.
- Autonomously prescribing antibiotics for confirmed strep throat may be much lower risk than autonomously starting a thrombolytic order set in a stroke workflow, even with the same level of clinician supervision.
Four refinements in this section matter most for practices.
A disclaimer doesn't lower the risk
FDA treats "directiveness" as a spectrum, not a yes-or-no question. The paper says an informational function doesn't become any less directive just because it includes a disclaimer. What counts is the substance and context of the output.
That undercuts a common vendor move: adding "this is not medical advice" to text that plainly tells the user what to do.
Patient-facing tools may count as higher risk
The paper notes that clinicians can usually spot a bad output and ignore it. Patients often can't, and they may act on it by delaying care or treating themselves. FDA suggests patient-facing functions could be moved higher on the consequences axis than the same function aimed at a clinician.
FDA also acknowledges the other side of this: democratizing clinical information, patient empowerment, and "avoiding medical parentalism."
Specialist-level tools used by generalists
FDA asks (Question 4) whether risk goes up when a tool designed around specialist knowledge is used by a clinician who lacks that specialty training. That is exactly how family physicians and nurse practitioners use AI for a dermatology or cardiology question.
Risk can change over a conversation
A chatbot might start by giving general information and, several turns later, be telling the user what to do. FDA says it is considering assessing risk across the whole conversation, not just one reply at a time.
For triage tools, FDA also wants both kinds of error counted. Under-escalation delays needed care. Over-escalation sends people to the ER unnecessarily, which causes anxiety, extra tests, and wasted resources, and can wear down patients' trust.
Idea 2: Test the AI the way you'd credential a clinician
This is the most novel part. FDA says testing methods built for software with fixed inputs and outputs "may not be appropriate" for generative AI, because the range of possible inputs is too large to test exhaustively.
Its proposed alternative is modeled on how human clinicians are evaluated: structured knowledge exams, then supervised practice, then ongoing assessment. The paper cites recent proposals from Patel and Blumenthal (JAMA Health Forum, 2026), Bergman, Wachter, and Emanuel (JAMA, 2026), and Freyer et al. (Nature Medicine, 2025).
The "competency-based approach" has two stages. Both would test the final product as deployed, not the underlying foundation model on its own.
Stage 1: Benchmarking
This is high-volume, non-clinical testing against a menu of elements. Not every element would apply to every device.
- Safety
- S.1: recognizing safety-critical situations and escalating them
- S.2: staying within the tool's intended scope
- S.3: communicating uncertainty honestly and deferring to a clinician when appropriate
- Clinical proficiency
- E.1: clinical knowledge
- E.2: gathering information and clinical reasoning
- E.3: calculations and measurements
- E.4: communication quality
- Generalizability
- R.1: robustness and reproducibility
- R.2: performance across patient subgroups
- Agentic capabilities
- A.1: planning multi-step tasks, using tools correctly, and stopping at human-oversight checkpoints
Appendix A describes each element in detail, and the descriptions are pointed. Under S.2, over-refusal (declining legitimate requests) counts as a failure alongside under-refusal. S.3 treats "presenting uncertain, outdated, or contested information with false confidence" as a safety failure. R.2 requires testing across "non-standard dialects, accents, colloquialisms, and lower general literacy and health literacy levels." E.4 names automation bias, where a user trusts an output because it sounds fluent or authoritative.
Expert graders would need to be independent of the vendor and of the foundation-model developer. FDA adds that this still applies "when the expert adjudicator is itself an LLM."
Stage 2: Clinical confirmation
FDA says clinical confirmation "might not require a prospective clinical study in every case." It lists options roughly in order of increasing rigor:
- Retrospective evaluation on real patient inputs
- Shadow deployment, where the tool runs on live patients but its outputs are hidden and don't affect care
- Standardized patient actors
- Clinician adjudication of real cases
- A prospective clinical study
The paper also raises a question with no easy answer: what should the AI be compared against? Options include a panel of clinicians reflecting the standard of care, a "median clinician in practice," the clinician-plus-AI team, or even what happens with no tool at all.
Idea 3: Accept more uncertainty up front, monitor more after launch
The paper asks outright (Question 18) whether FDA should "accept greater premarket uncertainty" about a generative AI device's benefits and risks in exchange for stronger postmarket monitoring. That is a real potential trade: faster market access in return for heavier ongoing surveillance.
Monitoring options FDA floats include:
- Periodic re-benchmarking, including after the underlying model changes
- Sample-based review of real-world outputs by independent clinicians
- Drift monitoring
- Possibly "machine-based supervisory agents," meaning AI monitoring AI
Two ideas deal with the fact that most of these products run on someone else's model:
- Change control for foundation-model updates. A foundation-model developer can change its model without the device maker doing anything. FDA asks (Question 24) how manufacturers can detect and respond to that, whether through contracts, technical measures, or a Predetermined Change Control Plan.
- Voluntary "Foundation Model Master Files." Model developers could confidentially give FDA structured model or system cards covering architecture, training data provenance, known failure modes, subgroup performance, guardrails, update-notification commitments, and audit logs. Device makers could then reference that file in their submissions. FDA is clear that filing one "would not constitute authorization of the underlying model." FDA also admits (Question 25) that model developers "may have limited incentive to disclose safety-relevant information."
The line to watch: who does the monitoring?
Under "Shared Ecosystem Responsibility," FDA lists clinicians, healthcare institutions, payers, professional societies, and state authorities as each having "potential roles to play" in monitoring and reporting. Question 21 asks how those roles can be structured "without diffusing manufacturer accountability."
That qualifier matters for practices. If postmarket monitoring replaces some premarket evidence, some of the monitoring work could shift to the people using the tools. A solo practice has no AI governance committee.
This echoes the EHR model-card story. Federal transparency requirements for AI inside certified EHRs may be cut back under HTI-5 at the same moment FDA is asking who should watch generative AI in the field.
Where the AI tools you actually use land
The paper doesn't name products, and nothing in it changes any tool's regulatory status today. Here is how its framework maps onto the categories in this directory.
AI scribes and documentation
The paper covers these directly, if briefly. In its section on agentic AI, it says such systems are being deployed for "care coordination, clinical documentation, patient outreach, and clinical workflow support, some or all of which may not be functions that are the focus of FDA's device regulatory oversight."
So for AI medical scribes and therapy-note tools, the practical answer is nothing new from FDA. That matches the January CDS guidance, which put clinician-reviewed summaries and drafted notes on the non-device side of the line.
Staying out of FDA's lane still doesn't mean anyone checked the tool. The paper's benchmarking list works as a free evaluation checklist anyway. R.2 (accents, dialects, low health literacy) and S.3 (false confidence) describe the same failure modes covered in what the research says about scribe accuracy.
Clinical decision support and clinical Q&A
Tools like Glass Health sit at the lower-left of FDA's chart when they give a clinician sourced, reviewable information. A footnote in the paper ties the "clinician can check the output" idea directly to the independent-review criterion of the CDS exclusion in section 520(o)(1)(E) of the FD&C Act.
The questions to watch are directiveness (how firmly the tool steers toward one answer) and the generalist-versus-specialist question. Both matter to a family physician or NP asking an AI about a condition outside their specialty. Our Glass Health review covers what that product claims today, and you can compare CDS tools for family medicine.
Patient-facing chat and voice agents
Scheduling, reminders, and intake, the core of tools like Hyro and Klara, are administrative work, not device functions. The voice-agent post covers the rules that actually apply there (TCPA, HIPAA, accessibility).
The paper's warning is about drift. A patient-facing bot that starts answering "should I come in?" is doing care-escalation triage. That is precisely the function FDA singles out, and patient-facing tools may sit higher on the risk chart. If your patient communication vendor is adding symptom questions, ask where the vendor thinks that feature lands.
Imaging
AI imaging and diagnostics tools such as Pearl and Overjet were already regulated devices and remain so. Question 17 asks whether the competency approach fits multimodal vision-language models, which is where generative AI and imaging start to overlap. Our dental AI guide covers the current clearances.
A real example: the first "patient-facing LLM" clearance
FDA's insulin example isn't hypothetical. On December 23, 2025, FDA cleared UpDoc (K253281), a prescription app for insulin management in adults with type 2 diabetes. It was cleared as a Class II drug-dose calculator (product code NDC, 21 CFR 868.1890), with Hygieia's d-Nav System as the predicate device.
UpDoc announced in June 2026 that this was the first FDA clearance for medical software using "patient-facing large language models."
The public record is narrower than that headline:
- The 510(k) summary describes a "Conversation Service" and a separate "Clinical Service." It does not use the term "large language model."
- The summary states that no clinical testing was performed. Clearance rested on non-clinical testing.
- The Predetermined Change Control Plan allows future changes only if they keep the insulin dosing logic deterministic.
In other words, the conversation layer collects and relays information, and the dosing decision is made by conventional, validated rules. STAT reported that the company's CEO wouldn't say whether the generative AI makes treatment decisions.
That is the architecture FDA's chart rewards: put the language model on the low-risk side of the chart and keep the high-stakes step deterministic. It is also a reminder to read the 510(k) summary, not the press release.
Should you comment? Probably yes, on a few questions
You don't need to answer all 26 questions. FDA explicitly welcomes partial responses. Clinicians see things vendors don't, and a few questions ask for exactly that:
- Question 3: patient-facing risk. If you have watched patients act on chatbot advice, that is useful evidence.
- Question 4: generalist versus specialist use. Primary care and NP perspectives are underrepresented in device policy.
- Question 6: over- versus under-escalation. Front-line clinicians see both kinds of error.
- Question 14: the comparator. Should AI match the standard of care, or a "median clinician in practice"?
- Question 21: who monitors. This is the place to say what small practices can and cannot realistically take on.
Comments go to docket FDA-2026-N-7874 on Regulations.gov, due October 19, 2026. Specific examples from your own practice will carry more weight than general opinions.
The throughline
In January, FDA loosened the rules for software that helps a clinician think. In August, it sketched how it might hold software to account when it acts more like a clinician: test it like one, watch it like one, and worry more when the person reading the output can't check it.
Nothing changes on October 20. But the ideas in this paper, especially the risk map, competency benchmarking, and the monitoring-for-evidence trade, are likely to shape whatever guidance comes next. They also make a good checklist for any AI tool you're evaluating, regulated or not.
Browse clinical decision support tools, AI scribes, and patient communication tools in our directory, or see how state rules fit alongside federal ones in our 2026 state healthcare AI laws roundup.
This article is for general information only and is not legal, regulatory, or medical advice. FDA discussion papers do not represent draft or final guidance and create no requirements; a product's regulatory status depends on its specific intended use and functions. Confirm any tool's regulatory status directly with the vendor and with qualified regulatory counsel, and check the FDA docket for the current comment deadline before submitting.
Related articles
FDA Just Moved the Line Between a Clinical Tool and a Regulated Device
In January 2026, the FDA rewrote its clinical decision support guidance and pushed more AI software outside its oversight. Single recommendations are now allowed, documentation tools got clearer footing, and generative AI got almost no framework at all. Here is what changed — and why 'no FDA clearance needed' tells you nothing about whether a tool works.
Glass Health Review 2026: An AI Scribe With a Differential Diagnosis Engine (Read the Terms First)
Glass Health pairs ambient scribing with a cited differential diagnosis for $0 to $200 a month. We cover the features, pricing and independent studies, plus what most reviews skip: ads on the two cheapest plans, individual terms that authorize model training on PHI, and a BAA you have to opt in to.
Washington Spent 2026 Trying to Erase State AI Laws. The Ones That Reach Your Scribe Were Never on the List.
Executive Order 14365 does not mention healthcare once. Meanwhile Rhode Island and Louisiana wrote the ambient AI scribe into statute, five states rewrote what a therapist may do with one, and eight states put humans back in front of AI claim denials. Here is every 2026 healthcare AI law that actually binds a practice — and why preemption is not coming to save anyone.
The AI Model Card in Your EHR: What HTI-5 Would Delete, and How to Read Yours First
Since January 2025, every certified EHR has had to publish a plain-language disclosure of what its built-in AI was trained on, how it was validated, and what it should never be used for — free, at a public link. ASTP/ONC has proposed deleting that requirement. Here is how to find yours before it goes, and why your standalone AI scribe was never covered by it in the first place.