Key Takeaways
- ·Existing governance frameworks were built for predictive models with discrete, measurable outputs. Generative AI produces open-ended content where errors look valid on their face and can propagate silently through the EHR.
- ·FDA's public list of AI-enabled medical devices had 1,614 entries when we checked it in September 2026, and it doesn't yet identify which devices use large language models. In August 2026 FDA asked for public feedback on how to regulate generative AI-enabled devices. FDA status tells you little about these newer tools.
- ·Agentic AI shifts the governance question from 'who reviews the recommendation' to 'who authorized the action.' That's a fundamentally different accountability problem.
- ·Shadow AI, scope creep, and prompt injection are risks that traditional clinical decision support governance wasn't built to address.
- ·Governance committees evaluating these tools need to expand their evaluation criteria to include hallucination rate, use-case boundary enforcement, and cross-system workflow traceability.
The short answer
Traditional clinical decision support governance was designed for tools with defined inputs, bounded outputs, and measurable accuracy metrics. Generative and agentic AI break those assumptions in different ways. Generative AI produces open-ended content where hallucinations are hard to detect and can enter the legal medical record before anyone catches them. Agentic AI takes autonomous action across systems rather than producing outputs for human review. Governing these tools requires expanding your evaluation criteria, enforcing use-case boundaries at deployment, building explicit shadow AI policies, and creating accountability structures that can trace multi-agent workflows. Your existing framework is the starting point.
What Traditional Governance Was Designed to Handle
Most health system AI governance frameworks were built around a specific model type: the predictive algorithm. A sepsis prediction score, a readmission risk model, an imaging classifier. These tools take defined inputs, produce measurable outputs, and can be evaluated using standard performance metrics like area under the curve, sensitivity, and specificity. The ground truth exists. The comparison is tractable.
That model type is still most of what FDA has reviewed. FDA's public list of AI-enabled medical devices had 1,614 entries when we checked it in September 2026. The list doesn't yet identify which devices use large language models. FDA says it "will explore methods to identify and tag" them in a future update. FDA's device framework grew up around algorithms that are locked, or that change in ways a manufacturer can specify in advance. In August 2026 the agency published a discussion paper on generative AI-enabled medical devices and asked for public feedback, with comments due October 19, 2026.
The practical implication for governance committees is significant. Health systems that have used FDA clearance as shorthand for "this tool has been vetted" shouldn't extend that logic to generative and agentic AI. Many generative tools in use today haven't been through FDA review at all, and whether a specific tool is a regulated device is a question for counsel. More of the vetting falls to your organization, and the frameworks that worked for traditional clinical decision support don't cover the new risk categories these tools introduce.
Where Generative AI Governance Diverges
The central governance challenge with generative AI is hallucination risk. These systems produce confident-sounding text that can be factually wrong, and unlike a predictive model that outputs a probability score you can evaluate against outcomes, a hallucinated clinical note or summary looks valid on its face. An AI scribe that generates a clinical note containing a medication error, or a patient instruction document that fabricates a follow-up recommendation, can enter the legal medical record and influence downstream decisions before any reviewer notices.
This changes what your validation metrics need to measure. Researchers at Duke published one practical example in 2025. Their SCRIBE framework for evaluating ambient scribing tools combines simulation, computational metrics, human reviewer assessment, and LLM-based evaluation. It scores generated notes on criteria such as fluency, completeness, and factuality, and adds simulations for robustness, bias, and fairness. In the authors' own test, the tool was strong on fluency and clarity and weaker on factual accuracy and on capturing new medications. Your governance committee needs to be equipped to evaluate these dimensions, which requires different expertise than reviewing an AUC score. A committee composed entirely of clinicians and administrators may not have the technical fluency to assess hallucination risk patterns, prompt sensitivity, or output coherence.
Three additional risk categories don't apply to traditional clinical decision support:
Scope creep
Tools deployed for one purpose can get repurposed for a different purpose without revalidation. A documentation scribe that starts being used as a differential diagnosis generator, or an ambient listening tool that clinicians begin treating as a clinical reference, is operating outside its validated use case. Governance must enforce use-case boundaries at the point of deployment and through ongoing monitoring, not only at initial approval. This requires both clear documentation of approved use cases and a mechanism to detect when actual use patterns diverge from them.
Shadow AI
Staff using consumer generative AI tools in clinical workflows can expose patient information in ways traditional clinical decision support never did. A physician summarizing a patient history in a consumer AI chatbot, a nurse using a free AI assistant to draft discharge instructions, a resident asking a general-purpose LLM to help interpret lab results. These behaviors call for explicit policy, staff training, and monitoring, which the CHAI Lifecycle Management playbook also recommends. Several AI vendors now offer enterprise or healthcare plans that include a business associate agreement, and free and personal accounts don't. Whether a particular use is permitted under HIPAA is a question for your privacy officer and counsel. The challenge is cultural as much as technical.
Workflow design for expected imperfection
Because hallucinations will occur, governance can't focus only on preventing errors. It has to operationalize human review in ways that catch them. The phrase 'human in the loop' needs to be a defined clinical process with specific standards for what reviewers are checking, how much time is allocated for that review, and what documentation is required when a reviewer accepts AI-generated content into the medical record. Vague commitments to oversight are hard to defend, and they don't catch errors.
Where Agentic AI Governance Diverges
Agentic AI is a different category of governance problem entirely. These systems go beyond supporting decisions and execute multi-step workflows: scheduling appointments based on care pathway logic, routing prior authorizations through payer systems, coordinating referrals across care teams, triggering communications to patients and providers. The governance question shifts from "who reviewed this recommendation" to "who authorized this action, and what happens when it's wrong."
Panelists at a HIMSS26 preconference session on AI operating models listed four things they considered essential for governing agents, starting with inventory: knowing which agents are operating in your environment, what data they can access, and what workflows they can initiate. From there, they move to data access controls that limit each agent's reach to what it needs for its authorized function, PHI and PII handling protocols that apply at each step of a multi-agent chain, and output monitoring that can surface when an agent has taken an action outside its defined parameters.
The accountability question cuts across all four. When multiple AI agents are passing work across your EHR, your scheduling system, and a payer interface, can you trace any specific action back to the agent that took it, the authorization that preceded it, and the data that informed it? If the answer is no, your governance framework doesn't yet have what you need for agentic systems.
Prompt injection and adversarial risk
Agentic systems that process external inputs, such as patient-submitted intake forms, incoming referral faxes, or data feeds from other vendors, are vulnerable to prompt injection attacks, where malicious or malformed content manipulates the agent into taking actions its operators didn't authorize. Traditional clinical decision support generally doesn't face this risk, because most of it doesn't process free-form external text. Agentic systems need input validation layers and sandboxed execution environments that your traditional governance framework never needed to contemplate.
Orchestration governance across vendors
When multiple agents from different vendors operate across the same systems, accountability for a given outcome can fall into the gap between them. Your governance framework needs to address cross-vendor handoffs explicitly: which vendor is accountable when an agent chain produces an error, how protocol adherence is monitored across systems, and what escalation triggers exist when an agent encounters a situation outside its defined parameters. Contracts with individual vendors don't solve this problem unless the orchestration question is addressed in each of them.
The Regulatory Gap for These Tool Categories
- ·FDA's public list of AI-enabled devices (1,614 entries in September 2026) doesn't yet identify which devices use large language models
- ·FDA's device framework developed around locked algorithms and around changes a manufacturer can specify in advance, which is the model its Predetermined Change Control Plan guidance follows
- ·FDA's Clinical Decision Support Software guidance (January 2026) addresses which decision support functions fall outside device regulation, so some tools in clinical use receive no FDA review
- ·In August 2026 FDA opened a request for feedback on regulating generative AI-enabled devices, with comments due October 19, 2026
- ·FDA status is a weak stand-in for local governance review of generative or agentic tools, so internal governance carries more weight than it did three years ago
Five Governance Differentiators in Practice
Translating the risk differences above into practical governance changes requires updating specific processes, not just adding language to a policy document.
Expand evaluation criteria
Add hallucination rate, output coherence, demographic fairness metrics, and use-case specificity to your governance committee's standard review template alongside traditional validation metrics. If your current template only asks for AUC and sensitivity, it's missing the dimensions that matter most for generative tools. The SCRIBE framework from Duke researchers is one published starting point.
Enforce use-case boundaries at deployment
Document the specific approved use cases for each generative tool at the time of governance approval. Build a process for flagging and reviewing use-case drift, because it will happen. A documentation tool that starts being used for differential diagnosis is the kind of drift to expect wherever use-case boundaries aren't enforced.
Build a shadow AI policy with teeth
Define which consumer AI tools are prohibited in clinical contexts, which are permitted with specific safeguards, and what training is required before any staff member uses a generative AI tool in a patient care workflow. The policy needs to address HIPAA exposure specifically, and the training needs to explain why the restrictions exist rather than just asserting them.
Operationalize human review
For each generative AI tool in clinical use, define specifically what reviewers are checking, how much time they're allocated to do it, what the escalation process is if they identify an error, and what documentation is required when they accept AI-generated content into the medical record. Vague commitments to oversight are hard to defend after an incident.
Require agentic traceability before deployment approval
For any agentic system, require architectural documentation showing how individual agent actions are logged, attributed, and auditable before your governance committee approves deployment. Retroactively adding traceability capability after an agentic system is in production is significantly harder than requiring it as a pre-deployment condition.
What a Governance Committee Needs to Do Differently
A multidisciplinary governance committee remains the right structure for AI oversight. But the composition and evaluation processes that worked for predictive analytics may not be adequate for generative and agentic tools. The expertise gaps are specific.
For generative AI, committee members need enough familiarity with LLM behavior to assess hallucination risk patterns, understand prompt sensitivity, and evaluate output coherence. This doesn't require deep technical expertise, but it does require more than traditional clinical informatics or quality improvement backgrounds typically provide. One option is adding a member with specific experience in AI model behavior or to bring in advisory expertise for high-risk tool reviews.
For agentic AI, the committee needs the ability to evaluate workflow architecture. Can they assess whether the agent's action boundaries are appropriately defined? Can they evaluate whether the logging and attribution mechanisms are adequate for incident investigation? Do they understand what cross-vendor handoff accountability looks like in practice? These are engineering and systems questions that governance committees weren't previously expected to engage with.
Traditional CDS governance needs
- ·Clinical informatics expertise
- ·Quality and patient safety representation
- ·Legal and compliance review
- ·IT and EHR integration assessment
- ·Standard metric interpretation (AUC, sensitivity, specificity)
Additional needs for generative and agentic AI
- ·LLM behavior and prompt sensitivity assessment
- ·Hallucination rate evaluation
- ·Shadow AI policy development
- ·Workflow architecture review for agentic systems
- ·Cross-vendor accountability framework design
Frequently Asked Questions
Common questions from health system leaders updating their AI governance frameworks for generative and agentic tools.
Is a hallucination from an AI clinical documentation tool a malpractice risk?
We aren't attorneys, and liability turns on the facts and on state law, so take this question to counsel. From a governance standpoint the concern is simple. A hallucinated note that influences a later care decision is a patient safety problem whether or not anyone is ever sued. A physician's sign-off doesn't remove the risk if the review behind it is too shallow to catch errors. That is why we recommend defining specific review standards in place of general human-in-the-loop language, and keeping a record that reviewers follow them.
What's the core governance difference between generative AI and traditional clinical decision support?
Traditional clinical decision support produces discrete, bounded outputs: a risk score, an alert, a recommendation from a defined set of options. Those outputs can be validated against outcomes, measured with standard metrics, and reviewed through established processes. Generative AI produces open-ended content where correctness is harder to define and errors are harder to detect at scale. The governance difference shows up in validation methods, monitoring approaches, and human review standards. All three need to be different for generative tools.
Does FDA review cover generative AI tools?
Sometimes, and you shouldn't assume it. A generative tool that meets the legal definition of a medical device can be subject to FDA oversight, and many tools used in health systems fall outside that definition or haven't been through FDA review. FDA's public list of AI-enabled devices doesn't yet identify which entries use large language models, and in August 2026 FDA asked for public feedback on how to regulate generative AI-enabled devices. Whether a specific tool is a regulated device is a question for counsel. Either way, FDA status doesn't replace your own governance review.
What is prompt injection and why does it matter for clinical AI?
Prompt injection is an attack where malicious or malformed input causes an AI system to behave in unintended ways. For agentic AI systems that process external text, such as patient intake forms, incoming referral documents, or data feeds from other vendors, a carefully crafted input could manipulate the agent into taking actions its operators didn't authorize. In a clinical context, that could mean unauthorized data access, incorrect workflow routing, or triggering inappropriate clinical communications. It's a risk category that traditional clinical decision support governance never had to address.
What is agentic AI and do we need to govern it differently from other AI tools?
Agentic AI refers to systems that take autonomous action rather than producing outputs for human review. Instead of recommending a care pathway, an agentic system might schedule the appointment, route the authorization, and send the coordination message without waiting for a human to act on each step. That requires governance addressing workflow boundaries, escalation triggers, cross-system accountability, and auditability in ways that traditional clinical decision support governance never contemplated. The core accountability question shifts from reviewing outputs to authorizing actions and tracing them after the fact.
How do we govern staff who are using ChatGPT or other consumer AI tools in their workflows?
Shadow AI governance requires three things: a clear policy that defines what's permitted and what's prohibited in clinical contexts, training that explains why the restrictions exist rather than just asserting them, and a monitoring approach that surfaces compliance issues without creating a punitive culture. The goal is to keep patient data out of systems that lack appropriate protections and agreements, and to make sure staff understand why. Several vendors now offer enterprise or healthcare plans that include a business associate agreement, which free and personal accounts don't. Which tools and uses are permitted for your organization is a decision to make with your privacy officer and counsel.
Sources
- U.S. Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices (public list). Content current as of September 4, 2026; 1,614 entries when accessed September 19, 2026.
- U.S. Food and Drug Administration. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. August 18, 2026. Docket FDA-2026-N-7874.
- U.S. Food and Drug Administration. Clinical Decision Support Software, final guidance. January 2026.
- U.S. Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions, final guidance.
- Wang H, Yang R, Alwakeel M, et al. An evaluation framework for ambient digital scribing tools in clinical applications. npj Digital Medicine 8, 358. Published June 13, 2025.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). July 26, 2024.
- The Joint Commission and the Coalition for Health AI. The Responsible Use of AI in Healthcare (RUAIH). September 17, 2025. (PDF)
- Coalition for Health AI. AI Governance Playbooks, Subdomain 4.1: Lifecycle Management. Released May 27, 2026.
- Vendor documentation on business associate agreements, cited as examples only: OpenAI for Healthcare (January 8, 2026) and Anthropic HIPAA-ready Enterprise plans. Mosaic Life Tech has no commercial relationship with either.
- "The New AI Operating Model for Healthcare," panel, AI in Healthcare Forum, HIMSS26, Las Vegas, March 9, 2026. Authors' notes from attendance.
Related Questions
- ›What should we require from AI vendors as a condition of deployment?
- ›How do we validate that an AI tool performs safely and equitably across our patient population?
- ›How do we create and maintain an AI tool inventory for our health system?
- ›What should our board be asking about AI risk?
- ›How to build an AI governance committee for a health system

