Healthcare AI Governance

    How Should Our Health System Govern Generative AI and Agentic AI Differently from Traditional Clinical Decision Support?

    The governance frameworks most health systems built for predictive analytics weren't designed for AI that generates open-ended clinical content or executes workflows autonomously. Here's where the gaps are and what to do about them.

    Last updated: · By Teresa Younkin & Jim Younkin, Mosaic Life Tech

    Key Takeaways

    • ·Existing governance frameworks were built for predictive models with discrete, measurable outputs. Generative AI produces open-ended content where errors look valid on their face and can propagate silently through the EHR.
    • ·FDA has cleared more than 1,200 AI-enabled medical devices but none with generative AI or LLM architecture. FDA clearance is not a reliable governance proxy for these newer tool categories.
    • ·Agentic AI shifts the governance question from 'who reviews the recommendation' to 'who authorized the action.' That's a fundamentally different accountability problem.
    • ·Shadow AI, scope creep, and prompt injection are risk categories unique to generative and agentic tools. Traditional clinical decision support governance doesn't address any of them.
    • ·Governance committees evaluating these tools need to expand their evaluation criteria to include hallucination rate, use-case boundary enforcement, and cross-system workflow traceability.

    The short answer

    Traditional clinical decision support governance was designed for tools with defined inputs, bounded outputs, and measurable accuracy metrics. Generative and agentic AI break those assumptions in different ways. Generative AI produces open-ended content where hallucinations are hard to detect and can enter the legal medical record before anyone catches them. Agentic AI takes autonomous action across systems rather than producing outputs for human review. Governing these tools requires expanding your evaluation criteria, enforcing use-case boundaries at deployment, building explicit shadow AI policies, and creating accountability structures that can trace multi-agent workflows. Your existing framework is a starting point, not a solution.

    What Traditional Governance Was Designed to Handle

    Most health system AI governance frameworks were built around a specific model type: the predictive algorithm. A sepsis prediction score, a readmission risk model, an imaging classifier. These tools take defined inputs, produce measurable outputs, and can be evaluated using standard performance metrics like area under the curve, sensitivity, and specificity. The ground truth exists. The comparison is tractable.

    That model type is still the majority of what the FDA has reviewed and cleared. Of the 1,200-plus AI-enabled medical devices cleared as of early 2026, none use generative AI capabilities or large language model architecture. The FDA's regulatory paradigm was built for adaptive medical devices with bounded behavior, not for systems that generate novel content or execute autonomous decisions across interconnected systems.

    The practical implication for governance committees is significant. Health systems that have historically used FDA clearance as shorthand for "this tool has been vetted" cannot extend that logic to generative and agentic AI. The vetting burden shifts substantially to your organization, and the frameworks that worked for traditional clinical decision support tools don't cover the new risk categories these tools introduce.

    Where Generative AI Governance Diverges

    The central governance challenge with generative AI is hallucination risk. These systems produce confident-sounding text that can be factually wrong, and unlike a predictive model that outputs a probability score you can evaluate against outcomes, a hallucinated clinical note or summary looks valid on its face. An AI scribe that generates a clinical note containing a medication error, or a patient instruction document that fabricates a follow-up recommendation, can enter the legal medical record and influence downstream decisions before any reviewer notices.

    This changes what your validation metrics need to measure. Duke Health's SCRIBE framework is one practical example of how leading health systems are adapting their evaluation criteria: expanding the standard template to include accuracy, fairness, and hallucination rate alongside the traditional predictive performance measures. Your governance committee needs to be equipped to evaluate these dimensions, which requires different expertise than reviewing an AUC score. A committee composed entirely of clinicians and administrators may not have the technical fluency to assess hallucination risk patterns, prompt sensitivity, or output coherence.

    Three additional risk categories don't apply to traditional clinical decision support:

    Scope creep

    Tools deployed for one purpose routinely get repurposed for a different purpose without revalidation. A documentation scribe that starts being used as a differential diagnosis generator, or an ambient listening tool that clinicians begin treating as a clinical reference, is operating outside its validated use case. Governance must enforce use-case boundaries at the point of deployment and through ongoing monitoring, not only at initial approval. This requires both clear documentation of approved use cases and a mechanism to detect when actual use patterns diverge from them.

    Shadow AI

    Staff using consumer generative AI tools in clinical workflows creates ungoverned HIPAA exposure that traditional clinical decision support never presented. A physician summarizing a patient history in a consumer AI chatbot, a nurse using a free AI assistant to draft discharge instructions, a resident asking a general-purpose LLM to help interpret lab results. These behaviors exist in most health systems today, and they require explicit policy, staff training, and monitoring structures that most governance frameworks don't address. The challenge isn't just technical. It's cultural.

    Workflow design for expected imperfection

    Because hallucinations will occur, governance can't focus only on preventing errors. It has to operationalize human review in ways that actually catch them. The phrase 'human in the loop' needs to be a defined clinical process with specific standards for what reviewers are checking, how much time is allocated for that review, and what documentation is required when a reviewer accepts AI-generated content into the medical record. Vague commitments to oversight don't hold up to scrutiny, and they don't catch errors.

    Where Agentic AI Governance Diverges

    Agentic AI is a different category of governance problem entirely. These systems don't just support decisions. They execute multi-step workflows: scheduling appointments based on care pathway logic, routing prior authorizations through payer systems, coordinating referrals across care teams, triggering communications to patients and providers. The governance question shifts from "who reviewed this recommendation" to "who authorized this action, and what happens when it's wrong."

    The four governance requirements identified for agentic AI at HIMSS 2026 start with inventory: knowing which agents are operating in your environment, what data they can access, and what workflows they can initiate. From there, they move to data access controls that limit each agent's reach to what it needs for its authorized function, PHI and PII handling protocols that apply at each step of a multi-agent chain, and output monitoring that can surface when an agent has taken an action outside its defined parameters.

    But the accountability question cuts across all four requirements. When multiple AI agents are passing work across your EHR, your scheduling system, and a payer interface, can you trace any specific action back to the agent that took it, the authorization that preceded it, and the data that informed it? If the answer is no, your governance framework doesn't yet have what you need for agentic systems.

    Prompt injection and adversarial risk

    Agentic systems that process external inputs, such as patient-submitted intake forms, incoming referral faxes, or data feeds from other vendors, are vulnerable to prompt injection attacks, where malicious or malformed content manipulates the agent into taking actions its operators didn't authorize. Traditional clinical decision support doesn't face this risk because it doesn't process free-form external text. Agentic systems need input validation layers and sandboxed execution environments that your traditional governance framework never needed to contemplate.

    Orchestration governance across vendors

    When multiple agents from different vendors operate across the same systems, accountability for a given outcome can fall into the gap between them. Your governance framework needs to address cross-vendor handoffs explicitly: which vendor is accountable when an agent chain produces an error, how protocol adherence is monitored across systems, and what escalation triggers exist when an agent encounters a situation outside its defined parameters. Contracts with individual vendors don't solve this problem unless the orchestration question is addressed in each of them.

    The Regulatory Gap for These Tool Categories

    • ·None of the 1,200-plus FDA-cleared AI medical devices use generative AI or LLM architecture as of early 2026
    • ·FDA's existing approval pathways were designed for bounded, adaptive AI, not for systems generating novel content
    • ·The Predetermined Change Control Plan (PCCP) framework allows algorithm updates without new submissions, but doesn't cover generative AI behavior changes
    • ·Post-market surveillance for AI remains underdeveloped, with no centralized reporting mechanism for AI performance failures
    • ·Health systems cannot rely on regulatory clearance as a governance proxy for generative or agentic tools, which means internal governance carries more weight than it did three years ago

    Five Governance Differentiators in Practice

    Translating the risk differences above into practical governance changes requires updating specific processes, not just adding language to a policy document.

    01

    Expand evaluation criteria

    Add hallucination rate, output coherence, demographic fairness metrics, and use-case specificity to your governance committee's standard review template alongside traditional validation metrics. If your current template only asks for AUC and sensitivity, it's missing the dimensions that matter most for generative tools. Duke's SCRIBE framework provides a practical starting point for what an expanded template looks like.

    02

    Enforce use-case boundaries at deployment

    Document the specific approved use cases for each generative tool at the time of governance approval. Build a process for flagging and reviewing use-case drift, because it will happen. A documentation tool that starts being used for differential diagnosis isn't a hypothetical risk. It's a pattern that has emerged at health systems without use-case boundary enforcement.

    03

    Build a shadow AI policy with teeth

    Define which consumer AI tools are prohibited in clinical contexts, which are permitted with specific safeguards, and what training is required before any staff member uses a generative AI tool in a patient care workflow. The policy needs to address HIPAA exposure specifically, and the training needs to explain why the restrictions exist rather than just asserting them. Compliance rates differ significantly between the two approaches.

    04

    Operationalize human review

    For each generative AI tool in clinical use, define specifically what reviewers are checking, how much time they're allocated to do it, what the escalation process is if they identify an error, and what documentation is required when they accept AI-generated content into the medical record. Vague commitments to oversight don't hold up in incident investigations.

    05

    Require agentic traceability before deployment approval

    For any agentic system, require architectural documentation showing how individual agent actions are logged, attributed, and auditable before your governance committee approves deployment. Retroactively adding traceability capability after an agentic system is in production is significantly harder than requiring it as a pre-deployment condition.

    What a Governance Committee Needs to Do Differently

    A multidisciplinary governance committee remains the right structure for AI oversight. But the composition and evaluation processes that worked for predictive analytics may not be adequate for generative and agentic tools. The expertise gaps are specific.

    For generative AI, committee members need enough familiarity with LLM behavior to assess hallucination risk patterns, understand prompt sensitivity, and evaluate output coherence. This doesn't require deep technical expertise, but it does require more than traditional clinical informatics or quality improvement backgrounds typically provide. Many committees are addressing this by adding a member with specific experience in AI model behavior or bringing in advisory expertise for high-risk tool reviews.

    For agentic AI, the committee needs the ability to evaluate workflow architecture. Can they assess whether the agent's action boundaries are appropriately defined? Can they evaluate whether the logging and attribution mechanisms are adequate for incident investigation? Do they understand what cross-vendor handoff accountability looks like in practice? These are engineering and systems questions that governance committees weren't previously expected to engage with.

    Traditional CDS governance needs

    • ·Clinical informatics expertise
    • ·Quality and patient safety representation
    • ·Legal and compliance review
    • ·IT and EHR integration assessment
    • ·Standard metric interpretation (AUC, sensitivity, specificity)

    Additional needs for generative and agentic AI

    • ·LLM behavior and prompt sensitivity assessment
    • ·Hallucination rate evaluation
    • ·Shadow AI policy development
    • ·Workflow architecture review for agentic systems
    • ·Cross-vendor accountability framework design

    Frequently Asked Questions

    Common questions from health system leaders updating their AI governance frameworks for generative and agentic tools.

    Is a hallucination from an AI clinical documentation tool a malpractice risk?

    It can be. If a hallucinated clinical note influences a subsequent care decision and that decision leads to patient harm, the liability question is whether the responsible clinician had an adequate review process in place. Courts and regulators will ask whether the health system's governance included meaningful human oversight of AI-generated content. The risk isn't eliminated by having a physician sign off on a note if the review process is too shallow to catch errors. This is exactly why governance frameworks need to define specific review standards rather than relying on general human-in-the-loop language.

    What's the core governance difference between generative AI and traditional clinical decision support?

    Traditional clinical decision support produces discrete, bounded outputs: a risk score, an alert, a recommendation from a defined set of options. Those outputs can be validated against outcomes, measured with standard metrics, and reviewed through established processes. Generative AI produces open-ended content where correctness is harder to define and errors are harder to detect at scale. The governance difference shows up in validation methods, monitoring approaches, and human review standards. All three need to be different for generative tools.

    Does FDA clearance apply to generative AI tools?

    No. As of early 2026, the FDA has cleared more than 1,200 AI-enabled medical devices, but none use generative AI capabilities or large language model architecture. The FDA's traditional regulatory framework was designed for adaptive AI with bounded, measurable outputs, not for systems that generate novel content. Health systems cannot use FDA clearance as a governance proxy for generative AI tools, regardless of any regulatory status a vendor claims. Local governance review is required.

    What is prompt injection and why does it matter for clinical AI?

    Prompt injection is an attack where malicious or malformed input causes an AI system to behave in unintended ways. For agentic AI systems that process external text, such as patient intake forms, incoming referral documents, or data feeds from other vendors, a carefully crafted input could manipulate the agent into taking actions its operators didn't authorize. In a clinical context, that could mean unauthorized data access, incorrect workflow routing, or triggering inappropriate clinical communications. It's a risk category that traditional clinical decision support governance never had to address.

    What is agentic AI and do we need to govern it differently from other AI tools?

    Agentic AI refers to systems that take autonomous action rather than producing outputs for human review. Instead of recommending a care pathway, an agentic system might schedule the appointment, route the authorization, and send the coordination message without waiting for a human to act on each step. That requires governance addressing workflow boundaries, escalation triggers, cross-system accountability, and auditability in ways that traditional clinical decision support governance never contemplated. The core accountability question shifts from reviewing outputs to authorizing actions and tracing them after the fact.

    How do we govern staff who are using ChatGPT or other consumer AI tools in their workflows?

    Shadow AI governance requires three things: a clear policy that defines what's permitted and what's prohibited in clinical contexts, training that explains why the restrictions exist rather than just asserting them, and a monitoring approach that surfaces compliance issues without creating a punitive culture. The goal isn't to prohibit all consumer AI use. It's to ensure that patient data isn't being processed through systems without appropriate data governance protections, and that staff understand the HIPAA exposure created when patient information enters a consumer AI platform.

    Sources

    • U.S. Food and Drug Administration. Artificial Intelligence and Machine Learning in Software as a Medical Device. Updated 2026.
    • Duke Health. SCRIBE Framework for Generative AI Evaluation: Accuracy, Fairness, and Coherence Metrics. 2025.
    • Rhew, D. "Four Non-Negotiable Elements of Agentic AI Governance." HIMSS Annual Conference. 2026.
    • Coalition for Health AI (CHAI). Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare. 2023.
    • The Joint Commission and CHAI. Responsible Use of Artificial Intelligence in Healthcare: Joint Guidance. 2024.
    • CHIME. AI Principles for Health Information and Technology. 2025.
    • NIST AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. January 2023.

    About the Authors

    Teresa Younkin

    Teresa Younkin, MSHI

    CEO & Co-Founder, Mosaic Life Tech

    20+ years leading AI, data governance, and interoperability initiatives across provider, payer, and federal health IT environments, including HL7 Da Vinci standards work and ONC programs.

    Jim Younkin

    Jim Younkin, MBA, FACHDM

    CTO & Co-Founder, Mosaic Life Tech

    30+ years across federal health IT programs, enterprise interoperability, and AI governance, including directing federal AI initiatives for ONC/ASTP and co-founding Pennsylvania's first regional HIE serving 4M+ patients.

    Mosaic Life Tech helps healthcare executives build board-visible AI governance posture aligned with Joint Commission and CHAI guidance. We don't sell AI tools or implementation services. Our work is advisory, and our interest is in helping organizations govern well before expectations harden into standards.

    Governing generative or agentic AI in your health system?

    We help health system leaders understand where their existing governance frameworks fall short for generative and agentic tools, and what updating them actually requires. Start with a conversation about what you're navigating.

    Start a Conversation