Healthcare AI Governance

    How Do We Validate That an AI Tool Performs Safely and Equitably Across Our Patient Population Before Go-Live?

    FDA clearance and vendor validation studies tell you how a tool performed somewhere else. Local validation tells you how it performs on your patients. The Joint Commission and CHAI guidance recommends looking at both.

    Last updated: · By Teresa Younkin & Jim Younkin, Mosaic Life Tech

    Key Takeaways

    • ·FDA clearance doesn't guarantee an AI tool performs safely on your specific patient population. A model trained on data from an academic medical center in Boston may not be appropriate for a rural clinic in Alabama.
    • ·Local validation means testing vendor-reported performance against your own patient population using your own data. At HIMSS26, the chair of Sutter Health's imaging service line described tracking two objective measures, cancer detection rate and callback rate, as its breast imaging AI rolled out.
    • ·Shadow deployment, running AI in silent mode alongside existing workflows before full go-live, is a widely used way to evaluate risk and bias before integration. The CHAI Lifecycle Management playbook describes silent evaluation as part of pre-deployment.
    • ·We recommend stratifying bias audits by race, age, sex, and socioeconomic status, using your local patient data in addition to what the vendor reports.
    • ·The Joint Commission and CHAI guidance says organizations should find out whether a tool was tested for the populations they serve and tuned or tested on local data. Separately, a federal rule, 45 CFR 92.210, calls for reasonable efforts to identify and mitigate discrimination risk from patient care decision support tools.
    • ·AI upgrades require re-validation. A new model version is not the same tool you approved.

    The short answer

    Local validation means running the AI against your own patient population, using your own data, before relying on it clinically. It starts with a shadow deployment phase where the AI generates outputs but they're not acted upon, which lets you evaluate accuracy and bias risk without affecting patient care. It includes bias audits stratified by the demographic characteristics of your specific patient mix, using your local data in addition to what the vendor reports. It produces documented results showing the gap between vendor-claimed performance and locally validated performance. And it establishes go/no-go criteria that your governance committee decided before validation began. How much local validation a tool needs should scale with its risk, a point the CHAI Lifecycle Management playbook makes directly.

    Why Vendor Validation Studies Aren't Enough

    When an AI vendor presents validation evidence, they're showing you how their tool performed in the conditions they chose to study. Those conditions, population demographics, clinical environment, workflow context, and data quality, are almost never identical to yours. A tool trained and validated at a large urban academic medical center may perform meaningfully differently when deployed at a community hospital serving a different patient mix. The Joint Commission and CHAI guidance gives a plain example: a tool developed on data from mostly younger, healthy patients may perform worse for older patients.

    Federal transparency rules help only partly. ONC's HTI-1 final rule (January 2024) requires developers of certified health IT to give users information about the predictive decision support tools they supply, so users can judge a tool's quality for themselves. It doesn't require providers to validate locally, and in December 2025 ONC proposed, in its HTI-5 rule, to remove those transparency requirements. FDA clearance is a separate matter. For devices that reach the market through the 510(k) pathway, clearance means FDA found the device substantially equivalent to one already legally marketed. That is a regulatory finding about the device, and it says nothing specific about how the tool performs on your patients.

    The guidance treats this as a question to raise at procurement. It says organizations should ask vendors how a tool was tested and validated for its intended use, whether the vendor is willing to tune or validate on a sample that represents the deployment setting, and how relevant biases were evaluated. An organization that relied entirely on vendor-supplied evidence hasn't answered those questions for itself.

    The Shadow Deployment Phase

    A shadow deployment, also called silent mode deployment, runs the AI in parallel with existing workflows so that outputs are generated and visible to evaluators, but not acted upon by clinicians. This phase lets you assess AI performance against real cases without any clinical consequences, and it allows comparison against current operational standards before the tool goes live.

    The shadow deployment phase serves several governance functions at once. It gives your clinical informatics team access to actual AI outputs across a representative sample of your patient population, which is the foundation for bias audits and performance analysis. It lets you identify whether the tool's output format and integration with your EHR workflows work as the vendor described. And it gives clinical staff familiarity with the tool's behavior before it influences care decisions, which improves the quality of the human review that follows go-live.

    Define the evaluation period before you start

    Establish how long the shadow deployment will run, what volume of cases is needed for adequate statistical power, and which patient subgroups need to be represented. Setting these parameters in advance prevents the evaluation from becoming open-ended or from stopping as soon as initial results look favorable.

    Establish go/no-go criteria before you see the results

    Your governance committee should define the performance thresholds that determine whether the tool proceeds to full deployment before validation begins. Defining criteria after you've already seen promising results, or after you've invested in a vendor relationship, creates pressure to rationalize rather than evaluate.

    Use the shadow phase to assess workflow integration

    Shadow deployment surfaces integration problems that weren't visible in vendor demos: whether outputs display correctly in your EHR, whether alert fatigue is a concern, whether the workflow assumptions built into the tool match how your clinicians work.

    Document the comparison against current standards

    Evaluate shadow deployment outputs against current clinical performance on the same cases where you have historical data. The question is whether the AI performs better than what you're already doing, and by how much, and whether that improvement holds across all patient subgroups or only in aggregate.

    Bias Audits: What They Need to Cover

    AI tools can perform well in aggregate while performing materially worse for specific patient subgroups. An imaging AI with strong overall sensitivity might have significantly lower sensitivity for patients of a particular age or body type. A sepsis prediction model might generate higher false positive rates for patients from certain demographic backgrounds. Aggregate accuracy metrics hide these disparities.

    The Joint Commission and CHAI guidance (RUAIH) addresses this in Element 6, Risk and Bias Assessment. It says organizations should determine whether algorithms "are tested for the specific populations they serve and ensure they are appropriately tuned and/or tested on local data." It adds that, in addition to vendor-reported information, organizations should check for bias when validating on local data and again after deployment, either internally or with their vendors. The guidance is voluntary and doesn't list demographic categories. The dimensions below are our recommendation.

    One binding federal rule sits next to that guidance. Under 45 CFR 92.210, a covered entity "must not discriminate on the basis of race, color, national origin, sex, age, or disability in its health programs or activities through the use of patient care decision support tools." The same rule sets "an ongoing duty to make reasonable efforts to identify" tools that use those characteristics as inputs, and to make reasonable efforts to mitigate the resulting risk. Those two duties applied from May 1, 2025. Ask your attorney how the rule applies to your organization and whether anything about its enforcement has changed.

    Practical tools for running these analyses include AI Fairness 360, an open-source toolkit first released by IBM Research, and Google's Fairness Indicators. Both are designed to measure whether AI performance differs systematically across demographic subgroups. The analysis requires access to outcome data that lets you compare AI outputs to known clinical results, which is why bias audits work best when conducted on historical cases where the correct answer is already documented.

    Bias Audit Dimensions to Evaluate

    • ·Race and ethnicity: performance parity across all racial groups represented in your patient population
    • ·Age: accuracy at the extremes of age distribution, particularly pediatric and elderly populations
    • ·Sex and gender: whether performance differs between male and female patients, and how the tool handles non-binary gender data
    • ·Socioeconomic status: whether outcomes proxy variables in the training data introduced bias against lower-income patients
    • ·Comorbidity patterns: whether the tool performs differently for patients with multiple conditions versus primary diagnoses
    • ·Geographic and facility type: if validating across multiple sites, whether performance is consistent or varies by location

    A Reference Point: Sutter Health's Objective Measures

    At HIMSS26 in March 2026, the chair of Sutter Health's imaging service line described how the system evaluated its breast imaging AI. Sutter didn't lean on clinicians' impressions of the tool. It tracked two objective measures, cancer detection rate and callback rate, through a pilot and then a wider rollout, and kept tracking them after go-live.

    When the tool moved from one major version to the next, Sutter tested the new version before relying on it. The principle applies broadly. A new model version can have different training data, different performance, and different bias patterns, so approval of one version shouldn't carry over to the next automatically. The Joint Commission and CHAI guidance makes a similar observation: AI tools and their underlying algorithms "may be updated periodically," which means performance can change.

    Ongoing measurement also addresses the performance drift problem. AI tools can degrade over time as patient populations shift, as data quality changes, or as clinical workflows evolve in ways that alter the distribution of inputs the AI receives. Continuous monitoring with objective metrics provides early warning of drift before it becomes a patient safety issue.

    Pre-deployment

    • ·Shadow deployment with objective KPIs
    • ·Bias audit stratified by local demographics
    • ·Go/no-go criteria set in advance
    • ·Gap analysis vs. vendor-claimed performance

    Go-live governance

    • ·Defined human review standards
    • ·Escalation process for AI errors
    • ·Integration monitoring for workflow issues
    • ·Initial performance reporting to governance committee

    Ongoing monitoring

    • ·Monthly or quarterly KPI review
    • ·Drift detection against baseline metrics
    • ·Re-validation triggered by model updates
    • ·Periodic bias re-audits as population shifts

    What to Document

    This is the record your governance committee and board should be able to produce:

    01

    Pre-deployment validation report

    The performance metrics from your shadow deployment phase, including overall accuracy against your patient population, results stratified by the demographic dimensions of your bias audit, the gap between vendor-claimed performance and locally measured performance, and the governance committee's go/no-go decision with rationale.

    02

    Bias audit results

    Documentation of the demographic subgroups evaluated, the metrics used to assess disparate impact, the specific results for each subgroup, and any remediation decisions made when disparities were identified. The audits should be conducted on your local population data, and the documentation should make that explicit.

    03

    Go-live criteria and approval record

    The performance thresholds your governance committee established before validation began, evidence that validation results met those thresholds, and the formal governance committee approval authorizing deployment. This record should predate go-live.

    04

    Ongoing monitoring reports

    Post-deployment performance data showing how the tool is performing against the KPIs established at go-live, including trend data showing whether performance is stable or drifting, and any exceptions or incidents that triggered governance review.

    05

    Re-validation records for model updates

    For each material vendor update to the AI model, documentation of the re-validation process, results, and governance committee approval before the update was deployed in your environment.

    The Health AI Partnership's Decision Points

    One of the Health AI Partnership's key decision points for adopting an AI tool is "Generate evidence of safety, efficacy and equity." Treat that as a real decision made before go-live, so deployment doesn't proceed by default because procurement has gone too far to reconsider. The governance committee decides on the evidence in front of it.

    Frequently Asked Questions

    Common questions from health system leaders building pre-deployment validation processes.

    What is local validation and how is it different from vendor validation?

    Local validation means testing an AI tool's performance against your own patient population, using your own data, before relying on it clinically. Vendor validation is testing the vendor conducted on the population they chose, under conditions they controlled. The two can differ. A tool validated at a large urban academic medical center may perform quite differently at a community hospital serving a different demographic mix. Local validation closes that gap by confirming the tool works for your patients.

    Does FDA clearance mean we don't need to validate locally?

    We recommend validating locally either way. For devices cleared through the 510(k) pathway, FDA describes clearance as a finding that the device is substantially equivalent to one already legally marketed. That is a regulatory finding about the device. It doesn't measure how the tool performs on your patient population. Treat FDA status as one input to your governance review.

    What is a shadow deployment and how do we run one?

    A shadow deployment, also called silent mode, runs an AI tool in parallel with existing workflows. The AI generates outputs that are visible to your evaluation team, but those outputs don't influence clinical decisions during the shadow phase. This lets you assess how the tool performs on real cases from your patient population without clinical consequences if the AI makes errors. Shadow deployment should run long enough to capture a representative sample of your patient mix and use cases, with performance evaluated against historical cases where the correct clinical answer is already documented.

    What does a bias audit for an AI tool involve?

    A bias audit assesses whether an AI tool performs differently across demographic subgroups. The audit compares AI performance on cases from different patient populations and asks whether accuracy, sensitivity, specificity, or false positive rates differ systematically by race, ethnicity, age, sex, or socioeconomic characteristics. We recommend running it on your local patient data, in addition to reviewing what the vendor reports, because demographic patterns in your patient mix may differ from those in the vendor's validation dataset. Open-source tools such as AI Fairness 360 support this analysis.

    Do we need to re-validate when a vendor updates their AI model?

    We recommend it. A new version can have different training data, different performance, and different bias patterns, so the approval that covered one version shouldn't automatically extend to the next. At HIMSS26, Sutter Health described testing a new major version of its breast imaging AI before relying on it. This is also why contracts need model change notification terms: you can't re-validate an update nobody told you about.

    What do the Joint Commission and CHAI say about validating AI tools?

    Their joint guidance (RUAIH, September 2025) is voluntary. It says organizations should ask vendors how a tool was tested and validated, find out whether it was tested for the populations they serve and tuned or tested on local data, check for bias before deployment and afterward, and keep monitoring performance on an ongoing, risk-based basis. On June 1, 2026 the Joint Commission began offering a voluntary certification based on that guidance, and accreditation isn't needed to apply. Organizations that can produce pre-deployment validation reports, bias audit results, and monitoring data are in a stronger position than those that can point only to vendor materials and FDA clearance.

    Sources

    About the Authors

    Teresa Younkin

    Teresa Younkin, MSHI

    CEO & Co-Founder, Mosaic Life Tech

    20+ years leading AI, data governance, and interoperability initiatives across provider, payer, and federal health IT environments, including HL7 Da Vinci standards work and ONC programs.

    Jim Younkin

    Jim Younkin, MBA, FACHDM

    CTO & Co-Founder, Mosaic Life Tech

    30+ years across federal health IT programs, enterprise interoperability, and AI governance, including directing federal AI initiatives for ONC and co-founding Pennsylvania's first regional HIE serving 4M+ patients.

    Mosaic Life Tech helps healthcare executives build board-visible AI governance posture in alignment with Joint Commission and CHAI guidance. We don't sell AI tools or represent vendors. Our work is advisory, we aren't attorneys, and we refer legal questions to counsel.

    Preparing to deploy an AI tool and not sure where to start on validation?

    We help healthcare executives build pre-deployment validation processes that produce the documentation governance committees and boards ask for. Start with a conversation about what you're working on.

    Start a Conversation