Healthcare AI Governance

    How Should Our Health System Govern Generative AI and Agentic AI Differently from Traditional Clinical Decision Support?

    The governance frameworks most health systems built for predictive analytics weren't designed for AI that generates open-ended clinical content or executes workflows autonomously. Here's where the gaps are and what to do about them.

    Last updated: · By Teresa Younkin & Jim Younkin, Mosaic Life Tech

    Key Takeaways

    • ·Existing governance frameworks were built for predictive models with discrete, measurable outputs. Generative AI produces open-ended content where errors look valid on their face and can propagate silently through the EHR.
    • ·FDA's public list of AI-enabled medical devices had 1,614 entries when we checked it in September 2026, and it doesn't yet identify which devices use large language models. In August 2026 FDA asked for public feedback on how to regulate generative AI-enabled devices. FDA status tells you little about these newer tools.
    • ·Agentic AI shifts the governance question from 'who reviews the recommendation' to 'who authorized the action.' That's a fundamentally different accountability problem.
    • ·Shadow AI, scope creep, and prompt injection are risks that traditional clinical decision support governance wasn't built to address.
    • ·Governance committees evaluating these tools need to expand their evaluation criteria to include hallucination rate, use-case boundary enforcement, and cross-system workflow traceability.

    The short answer

    Traditional clinical decision support governance was designed for tools with defined inputs, bounded outputs, and measurable accuracy metrics. Generative and agentic AI break those assumptions in different ways. Generative AI produces open-ended content where hallucinations are hard to detect and can enter the legal medical record before anyone catches them. Agentic AI takes autonomous action across systems rather than producing outputs for human review. Governing these tools requires expanding your evaluation criteria, enforcing use-case boundaries at deployment, building explicit shadow AI policies, and creating accountability structures that can trace multi-agent workflows. Your existing framework is the starting point.

    What Traditional Governance Was Designed to Handle

    Most health system AI governance frameworks were built around a specific model type: the predictive algorithm. A sepsis prediction score, a readmission risk model, an imaging classifier. These tools take defined inputs, produce measurable outputs, and can be evaluated using standard performance metrics like area under the curve, sensitivity, and specificity. The ground truth exists. The comparison is tractable.

    That model type is still most of what FDA has reviewed. FDA's public list of AI-enabled medical devices had 1,614 entries when we checked it in September 2026. The list doesn't yet identify which devices use large language models. FDA says it "will explore methods to identify and tag" them in a future update. FDA's device framework grew up around algorithms that are locked, or that change in ways a manufacturer can specify in advance. In August 2026 the agency published a discussion paper on generative AI-enabled medical devices and asked for public feedback, with comments due October 19, 2026.

    The practical implication for governance committees is significant. Health systems that have used FDA clearance as shorthand for "this tool has been vetted" shouldn't extend that logic to generative and agentic AI. Many generative tools in use today haven't been through FDA review at all, and whether a specific tool is a regulated device is a question for counsel. More of the vetting falls to your organization, and the frameworks that worked for traditional clinical decision support don't cover the new risk categories these tools introduce.

    Where Generative AI Governance Diverges

    The central governance challenge with generative AI is hallucination risk. These systems produce confident-sounding text that can be factually wrong, and unlike a predictive model that outputs a probability score you can evaluate against outcomes, a hallucinated clinical note or summary looks valid on its face. An AI scribe that generates a clinical note containing a medication error, or a patient instruction document that fabricates a follow-up recommendation, can enter the legal medical record and influence downstream decisions before any reviewer notices.

    This changes what your validation metrics need to measure. Researchers at Duke published one practical example in 2025. Their SCRIBE framework for evaluating ambient scribing tools combines simulation, computational metrics, human reviewer assessment, and LLM-based evaluation. It scores generated notes on criteria such as fluency, completeness, and factuality, and adds simulations for robustness, bias, and fairness. In the authors' own test, the tool was strong on fluency and clarity and weaker on factual accuracy and on capturing new medications. Your governance committee needs to be equipped to evaluate these dimensions, which requires different expertise than reviewing an AUC score. A committee composed entirely of clinicians and administrators may not have the technical fluency to assess hallucination risk patterns, prompt sensitivity, or output coherence.

    Three additional risk categories don't apply to traditional clinical decision support:

    Scope creep

    Tools deployed for one purpose can get repurposed for a different purpose without revalidation. A documentation scribe that starts being used as a differential diagnosis generator, or an ambient listening tool that clinicians begin treating as a clinical reference, is operating outside its validated use case. Governance must enforce use-case boundaries at the point of deployment and through ongoing monitoring, not only at initial approval. This requires both clear documentation of approved use cases and a mechanism to detect when actual use patterns diverge from them.

    Shadow AI

    Staff using consumer generative AI tools in clinical workflows can expose patient information in ways traditional clinical decision support never did. A physician summarizing a patient history in a consumer AI chatbot, a nurse using a free AI assistant to draft discharge instructions, a resident asking a general-purpose LLM to help interpret lab results. These behaviors call for explicit policy, staff training, and monitoring, which the CHAI Lifecycle Management playbook also recommends. Several AI vendors now offer enterprise or healthcare plans that include a business associate agreement, and free and personal accounts don't. Whether a particular use is permitted under HIPAA is a question for your privacy officer and counsel. The challenge is cultural as much as technical.

    Workflow design for expected imperfection

    Because hallucinations will occur, governance can't focus only on preventing errors. It has to operationalize human review in ways that catch them. The phrase 'human in the loop' needs to be a defined clinical process with specific standards for what reviewers are checking, how much time is allocated for that review, and what documentation is required when a reviewer accepts AI-generated content into the medical record. Vague commitments to oversight are hard to defend, and they don't catch errors.

    Where Agentic AI Governance Diverges

    Agentic AI is a different category of governance problem entirely. These systems go beyond supporting decisions and execute multi-step workflows: scheduling appointments based on care pathway logic, routing prior authorizations through payer systems, coordinating referrals across care teams, triggering communications to patients and providers. The governance question shifts from "who reviewed this recommendation" to "who authorized this action, and what happens when it's wrong."

    Panelists at a HIMSS26 preconference session on AI operating models listed four things they considered essential for governing agents, starting with inventory: knowing which agents are operating in your environment, what data they can access, and what workflows they can initiate. From there, they move to data access controls that limit each agent's reach to what it needs for its authorized function, PHI and PII handling protocols that apply at each step of a multi-agent chain, and output monitoring that can surface when an agent has taken an action outside its defined parameters.

    The accountability question cuts across all four. When multiple AI agents are passing work across your EHR, your scheduling system, and a payer interface, can you trace any specific action back to the agent that took it, the authorization that preceded it, and the data that informed it? If the answer is no, your governance framework doesn't yet have what you need for agentic systems.

    Prompt injection and adversarial risk

    Agentic systems that process external inputs, such as patient-submitted intake forms, incoming referral faxes, or data feeds from other vendors, are vulnerable to prompt injection attacks, where malicious or malformed content manipulates the agent into taking actions its operators didn't authorize. Traditional clinical decision support generally doesn't face this risk, because most of it doesn't process free-form external text. Agentic systems need input validation layers and sandboxed execution environments that your traditional governance framework never needed to contemplate.

    Orchestration governance across vendors

    When multiple agents from different vendors operate across the same systems, accountability for a given outcome can fall into the gap between them. Your governance framework needs to address cross-vendor handoffs explicitly: which vendor is accountable when an agent chain produces an error, how protocol adherence is monitored across systems, and what escalation triggers exist when an agent encounters a situation outside its defined parameters. Contracts with individual vendors don't solve this problem unless the orchestration question is addressed in each of them.

    The Regulatory Gap for These Tool Categories

    • ·FDA's public list of AI-enabled devices (1,614 entries in September 2026) doesn't yet identify which devices use large language models
    • ·FDA's device framework developed around locked algorithms and around changes a manufacturer can specify in advance, which is the model its Predetermined Change Control Plan guidance follows
    • ·FDA's Clinical Decision Support Software guidance (January 2026) addresses which decision support functions fall outside device regulation, so some tools in clinical use receive no FDA review
    • ·In August 2026 FDA opened a request for feedback on regulating generative AI-enabled devices, with comments due October 19, 2026
    • ·FDA status is a weak stand-in for local governance review of generative or agentic tools, so internal governance carries more weight than it did three years ago

    Five Governance Differentiators in Practice

    Translating the risk differences above into practical governance changes requires updating specific processes, not just adding language to a policy document.

    01

    Expand evaluation criteria

    Add hallucination rate, output coherence, demographic fairness metrics, and use-case specificity to your governance committee's standard review template alongside traditional validation metrics. If your current template only asks for AUC and sensitivity, it's missing the dimensions that matter most for generative tools. The SCRIBE framework from Duke researchers is one published starting point.

    02

    Enforce use-case boundaries at deployment

    Document the specific approved use cases for each generative tool at the time of governance approval. Build a process for flagging and reviewing use-case drift, because it will happen. A documentation tool that starts being used for differential diagnosis is the kind of drift to expect wherever use-case boundaries aren't enforced.

    03

    Build a shadow AI policy with teeth

    Define which consumer AI tools are prohibited in clinical contexts, which are permitted with specific safeguards, and what training is required before any staff member uses a generative AI tool in a patient care workflow. The policy needs to address HIPAA exposure specifically, and the training needs to explain why the restrictions exist rather than just asserting them.

    04

    Operationalize human review

    For each generative AI tool in clinical use, define specifically what reviewers are checking, how much time they're allocated to do it, what the escalation process is if they identify an error, and what documentation is required when they accept AI-generated content into the medical record. Vague commitments to oversight are hard to defend after an incident.

    05

    Require agentic traceability before deployment approval

    For any agentic system, require architectural documentation showing how individual agent actions are logged, attributed, and auditable before your governance committee approves deployment. Retroactively adding traceability capability after an agentic system is in production is significantly harder than requiring it as a pre-deployment condition.

    What a Governance Committee Needs to Do Differently

    A multidisciplinary governance committee remains the right structure for AI oversight. But the composition and evaluation processes that worked for predictive analytics may not be adequate for generative and agentic tools. The expertise gaps are specific.

    For generative AI, committee members need enough familiarity with LLM behavior to assess hallucination risk patterns, understand prompt sensitivity, and evaluate output coherence. This doesn't require deep technical expertise, but it does require more than traditional clinical informatics or quality improvement backgrounds typically provide. One option is adding a member with specific experience in AI model behavior or to bring in advisory expertise for high-risk tool reviews.

    For agentic AI, the committee needs the ability to evaluate workflow architecture. Can they assess whether the agent's action boundaries are appropriately defined? Can they evaluate whether the logging and attribution mechanisms are adequate for incident investigation? Do they understand what cross-vendor handoff accountability looks like in practice? These are engineering and systems questions that governance committees weren't previously expected to engage with.

    Traditional CDS governance needs

    • ·Clinical informatics expertise
    • ·Quality and patient safety representation
    • ·Legal and compliance review
    • ·IT and EHR integration assessment
    • ·Standard metric interpretation (AUC, sensitivity, specificity)

    Additional needs for generative and agentic AI

    • ·LLM behavior and prompt sensitivity assessment
    • ·Hallucination rate evaluation
    • ·Shadow AI policy development
    • ·Workflow architecture review for agentic systems
    • ·Cross-vendor accountability framework design

    Frequently Asked Questions

    Common questions from health system leaders updating their AI governance frameworks for generative and agentic tools.

    Is a hallucination from an AI clinical documentation tool a malpractice risk?

    We aren't attorneys, and liability turns on the facts and on state law, so take this question to counsel. From a governance standpoint the concern is simple. A hallucinated note that influences a later care decision is a patient safety problem whether or not anyone is ever sued. A physician's sign-off doesn't remove the risk if the review behind it is too shallow to catch errors. That is why we recommend defining specific review standards in place of general human-in-the-loop language, and keeping a record that reviewers follow them.

    What's the core governance difference between generative AI and traditional clinical decision support?

    Traditional clinical decision support produces discrete, bounded outputs: a risk score, an alert, a recommendation from a defined set of options. Those outputs can be validated against outcomes, measured with standard metrics, and reviewed through established processes. Generative AI produces open-ended content where correctness is harder to define and errors are harder to detect at scale. The governance difference shows up in validation methods, monitoring approaches, and human review standards. All three need to be different for generative tools.

    Does FDA review cover generative AI tools?

    Sometimes, and you shouldn't assume it. A generative tool that meets the legal definition of a medical device can be subject to FDA oversight, and many tools used in health systems fall outside that definition or haven't been through FDA review. FDA's public list of AI-enabled devices doesn't yet identify which entries use large language models, and in August 2026 FDA asked for public feedback on how to regulate generative AI-enabled devices. Whether a specific tool is a regulated device is a question for counsel. Either way, FDA status doesn't replace your own governance review.

    What is prompt injection and why does it matter for clinical AI?

    Prompt injection is an attack where malicious or malformed input causes an AI system to behave in unintended ways. For agentic AI systems that process external text, such as patient intake forms, incoming referral documents, or data feeds from other vendors, a carefully crafted input could manipulate the agent into taking actions its operators didn't authorize. In a clinical context, that could mean unauthorized data access, incorrect workflow routing, or triggering inappropriate clinical communications. It's a risk category that traditional clinical decision support governance never had to address.

    What is agentic AI and do we need to govern it differently from other AI tools?

    Agentic AI refers to systems that take autonomous action rather than producing outputs for human review. Instead of recommending a care pathway, an agentic system might schedule the appointment, route the authorization, and send the coordination message without waiting for a human to act on each step. That requires governance addressing workflow boundaries, escalation triggers, cross-system accountability, and auditability in ways that traditional clinical decision support governance never contemplated. The core accountability question shifts from reviewing outputs to authorizing actions and tracing them after the fact.

    How do we govern staff who are using ChatGPT or other consumer AI tools in their workflows?

    Shadow AI governance requires three things: a clear policy that defines what's permitted and what's prohibited in clinical contexts, training that explains why the restrictions exist rather than just asserting them, and a monitoring approach that surfaces compliance issues without creating a punitive culture. The goal is to keep patient data out of systems that lack appropriate protections and agreements, and to make sure staff understand why. Several vendors now offer enterprise or healthcare plans that include a business associate agreement, which free and personal accounts don't. Which tools and uses are permitted for your organization is a decision to make with your privacy officer and counsel.

    Sources

    About the Authors

    Teresa Younkin

    Teresa Younkin, MSHI

    CEO & Co-Founder, Mosaic Life Tech

    20+ years leading AI, data governance, and interoperability initiatives across provider, payer, and federal health IT environments, including HL7 Da Vinci standards work and ONC programs.

    Jim Younkin

    Jim Younkin, MBA, FACHDM

    CTO & Co-Founder, Mosaic Life Tech

    30+ years across federal health IT programs, enterprise interoperability, and AI governance, including directing federal AI initiatives for ONC and co-founding Pennsylvania's first regional HIE serving 4M+ patients.

    Mosaic Life Tech helps healthcare executives build board-visible AI governance posture in alignment with Joint Commission and CHAI guidance. We don't sell AI tools or represent vendors. Our work is advisory, we aren't attorneys, and we refer legal questions to counsel.

    Governing generative or agentic AI in your health system?

    We help health system leaders understand where their existing governance frameworks fall short for generative and agentic tools, and what updating them takes. Start with a conversation about what you're navigating.

    Start a Conversation