Key Takeaways
- ·FDA clearance doesn't guarantee an AI tool performs safely on your specific patient population. A model trained on data from an academic medical center in Boston may not be appropriate for a rural clinic in Alabama.
- ·Local validation means testing vendor performance against your own patient population using your own data. Sutter Health's approach, re-validating with objective KPIs monthly across their 26-hospital system, represents the practical standard.
- ·Shadow deployment, running AI in silent mode alongside existing workflows before full go-live, is now a recommended standard for evaluating risk and bias prior to integration.
- ·Bias audits must be stratified by race, age, gender, and socioeconomic status and conducted on your local patient population, not vendor-supplied data.
- ·Joint Commission surveyors are actively asking how organizations validated AI tools on their patient populations. Post-deployment validation reports are among the documentation surveyors look for.
- ·AI upgrades require re-validation. A new model version is not the same tool you approved.
The short answer
Local validation means running the AI against your own patient population, using your own data, before relying on it clinically. It starts with a shadow deployment phase where the AI generates outputs but they're not acted upon, which lets you evaluate accuracy and bias risk without affecting patient care. It includes bias audits stratified by the demographic characteristics of your specific patient mix, using your local data rather than vendor-supplied populations. It produces documented results showing the gap between vendor-claimed performance and locally validated performance. And it establishes go/no-go criteria that your governance committee decided before validation began, not after you've already committed to a vendor relationship.
Why Vendor Validation Studies Aren't Enough
When an AI vendor presents validation evidence, they're showing you how their tool performed in the conditions they chose to study. Those conditions, population demographics, clinical environment, workflow context, and data quality, are almost never identical to yours. A tool trained and validated at a large urban academic medical center may perform meaningfully differently when deployed at a community hospital serving a different patient mix. That gap is well documented and consistently underestimated by health systems during procurement.
The HTI-1 "Appropriateness" principle makes this explicit: a model trained on data from one setting may not be appropriate for another, and governance must enforce local validation to confirm that an AI tool actually works for your patients, not just for the population the vendor used for development. FDA clearance reinforces this point in a different way. Clearance means the FDA determined a device is substantially equivalent to a predicate device. It's a regulatory threshold, not a clinical performance guarantee for your specific patient population, and FDA-cleared AI tools have been suspended, recalled, and discontinued by health systems after post-deployment validation revealed performance problems.
The practical accountability implication is significant. Joint Commission surveyors are already asking health systems directly: how did you validate this tool on your patient population? Did you rely solely on vendor data or did you conduct your own evaluation? Organizations that relied entirely on vendor-supplied validation evidence are not in a strong position to answer that question.
The Shadow Deployment Phase
A shadow deployment, also called silent mode deployment, runs the AI in parallel with existing workflows so that outputs are generated and visible to evaluators, but not acted upon by clinicians. This phase lets you assess AI performance against real cases without any clinical consequences, and it allows comparison against current operational standards before the tool goes live.
The shadow deployment phase serves several governance functions at once. It gives your clinical informatics team access to actual AI outputs across a representative sample of your patient population, which is the foundation for bias audits and performance analysis. It lets you identify whether the tool's output format and integration with your EHR workflows actually work as the vendor described. And it gives clinical staff familiarity with the tool's behavior before it influences care decisions, which improves the quality of the human review that follows go-live.
Define the evaluation period before you start
Establish how long the shadow deployment will run, what volume of cases is needed for adequate statistical power, and which patient subgroups need to be represented. Setting these parameters in advance prevents the evaluation from becoming open-ended or from stopping as soon as initial results look favorable.
Establish go/no-go criteria before you see the results
Your governance committee should define the performance thresholds that determine whether the tool proceeds to full deployment before validation begins. Defining criteria after you've already seen promising results, or after you've invested in a vendor relationship, creates pressure to rationalize rather than evaluate.
Use the shadow phase to assess workflow integration
Shadow deployment surfaces integration problems that weren't visible in vendor demos: whether outputs display correctly in your EHR, whether alert fatigue is a concern, whether the workflow assumptions built into the tool match how your clinicians actually work.
Document the comparison against current standards
Evaluate shadow deployment outputs against current clinical performance on the same cases where you have historical data. The question is whether the AI performs better than what you're already doing, and by how much, and whether that improvement holds across all patient subgroups or only in aggregate.
Bias Audits: What They Need to Cover
AI tools can perform well in aggregate while performing materially worse for specific patient subgroups. An imaging AI with strong overall sensitivity might have significantly lower sensitivity for patients of a particular age or body type. A sepsis prediction model might generate higher false positive rates for patients from certain demographic backgrounds. Aggregate accuracy metrics hide these disparities.
The RUAIH framework, developed jointly by the Joint Commission and CHAI, addresses this directly. Element 6 of the framework requires proactive evaluation for disparate impact across racial groups, age, gender, and socioeconomic populations, and it specifies that assessments must be conducted on your local patient population, not vendor-supplied data. The distinction matters because demographic patterns in your patient population may differ significantly from those in the vendor's validation dataset.
Practical tools for running these analyses include AI Fairness 360, developed by IBM Research and widely used in healthcare contexts, and Google Fairness Indicators. Both are designed to measure whether AI performance differs systematically across demographic subgroups. The analysis requires access to outcome data that lets you compare AI outputs to known clinical results, which is why bias audits work best when conducted on historical cases where the correct answer is already documented.
Bias Audit Dimensions to Evaluate
- ·Race and ethnicity: performance parity across all racial groups represented in your patient population
- ·Age: accuracy at the extremes of age distribution, particularly pediatric and elderly populations
- ·Sex and gender: whether performance differs between male and female patients, and how the tool handles non-binary gender data
- ·Socioeconomic status: whether outcomes proxy variables in the training data introduced bias against lower-income patients
- ·Comorbidity patterns: whether the tool performs differently for patients with multiple conditions versus primary diagnoses
- ·Geographic and facility type: if validating across multiple sites, whether performance is consistent or varies by location
The Sutter Health Model: Objective KPIs, Continuous Monitoring
Sutter Health's approach to AI validation across their 26-hospital system provides a practical reference point. Rather than relying on subjective clinician impressions of AI performance, Sutter established objective KPIs measured on a monthly basis. For their breast cancer imaging AI, those metrics included cancer detection rate and callback rate, measured consistently across their system.
When their breast cancer AI vendor released an upgraded version, Sutter re-validated the new version before deploying it. This reflects a governance principle that leading health systems have internalized: an AI upgrade is not the same tool you approved. A new model version has different training data, different performance characteristics, and potentially different bias patterns. Approving version 1.0 is not approval of version 2.0.
The monthly monitoring cadence also addresses the performance drift problem. AI tools can degrade over time as patient populations shift, as data quality changes, or as clinical workflows evolve in ways that alter the distribution of inputs the AI receives. Continuous monitoring with objective metrics provides early warning of drift before it becomes a patient safety issue.
Pre-deployment
- ·Shadow deployment with objective KPIs
- ·Bias audit stratified by local demographics
- ·Go/no-go criteria set in advance
- ·Gap analysis vs. vendor-claimed performance
Go-live governance
- ·Defined human review standards
- ·Escalation process for AI errors
- ·Integration monitoring for workflow issues
- ·Initial performance reporting to governance committee
Ongoing monitoring
- ·Monthly or quarterly KPI review
- ·Drift detection against baseline metrics
- ·Re-validation triggered by model updates
- ·Periodic bias re-audits as population shifts
What to Document for Governance and Surveyor Review
Joint Commission surveyors probing AI governance are asking specific questions about validation. The documentation your governance program should be able to produce:
Pre-deployment validation report
The performance metrics from your shadow deployment phase, including overall accuracy against your patient population, results stratified by the demographic dimensions of your bias audit, the gap between vendor-claimed performance and locally measured performance, and the governance committee's go/no-go decision with rationale.
Bias audit results
Documentation of the demographic subgroups evaluated, the metrics used to assess disparate impact, the specific results for each subgroup, and any remediation decisions made when disparities were identified. The audits should be conducted on your local population data, and the documentation should make that explicit.
Go-live criteria and approval record
The performance thresholds your governance committee established before validation began, evidence that validation results met those thresholds, and the formal governance committee approval authorizing deployment. This record should predate go-live, not be reconstructed after the fact.
Ongoing monitoring reports
Post-deployment performance data showing how the tool is performing against the KPIs established at go-live, including trend data showing whether performance is stable or drifting, and any exceptions or incidents that triggered governance review.
Re-validation records for model updates
For each material vendor update to the AI model, documentation of the re-validation process, results, and governance committee approval before the update was deployed in your environment.
The HAIP Lifecycle Decision Gate
The Health AI Partnership's lifecycle guidance structures AI deployment around explicit decision gates, with the core question being: "Generate evidence of safety and efficacy, then determine if the AI should be integrated." This framing forces a hard decision before go-live rather than allowing deployment to proceed by default when procurement has advanced too far to reconsider. Your validation process should be structured around the same logic: the governance committee is making a binary decision based on evidence, not endorsing a procurement decision after the fact.
Frequently Asked Questions
Common questions from health system leaders building pre-deployment validation processes.
What is local validation and how is it different from vendor validation?
Local validation means testing an AI tool's performance against your own patient population, using your own data, before relying on it clinically. Vendor validation is testing the vendor conducted on the population they chose, under conditions they controlled. The gap between the two is often significant. A tool validated at a large urban academic medical center may perform quite differently at a community hospital serving a different demographic mix. Local validation closes that gap by confirming the tool works for your patients, not just for the vendor's test population.
Does FDA clearance mean we don't need to validate locally?
No. FDA clearance means the FDA determined a device is substantially equivalent to a predicate device. It's a regulatory threshold that establishes baseline safety and effectiveness, not a guarantee of performance in your specific patient population. Multiple FDA-cleared AI tools have been recalled or discontinued after post-deployment validation revealed performance problems in real clinical environments. FDA clearance is one input to your governance review, not a substitute for local validation.
What is a shadow deployment and how do we run one?
A shadow deployment, also called silent mode, runs an AI tool in parallel with existing workflows. The AI generates outputs that are visible to your evaluation team, but those outputs don't influence clinical decisions during the shadow phase. This lets you assess how the tool performs on real cases from your patient population without clinical consequences if the AI makes errors. Shadow deployment should run long enough to capture a representative sample of your patient mix and use cases, with performance evaluated against historical cases where the correct clinical answer is already documented.
What does a bias audit for an AI tool actually involve?
A bias audit assesses whether an AI tool performs differently across demographic subgroups. The audit compares AI performance on cases from different patient populations and asks whether accuracy, sensitivity, specificity, or false positive rates differ systematically by race, ethnicity, age, sex, or socioeconomic characteristics. The audit must be conducted on your local patient population data, not vendor-supplied data, because demographic patterns in your patient mix may differ from those in the vendor's validation dataset. Tools like IBM's AI Fairness 360 support this analysis.
Do we need to re-validate when a vendor updates their AI model?
Yes. An AI model update is not the same tool you validated. A new version has different training data, potentially different performance characteristics, and possibly different bias patterns. The governance approval that covered version 1.0 doesn't extend to version 2.0. Sutter Health's practice of re-validating before deploying new model versions, rather than assuming continuity of performance, reflects a governance standard that health systems should apply broadly. This is also why contracts need model change notification requirements: you can't re-validate a model update you weren't notified was happening.
What will a Joint Commission surveyor ask about our AI validation process?
Joint Commission surveyors are currently probing whether organizations validated AI tools on their own patient populations or relied solely on vendor data. Specific questions include how the organization assessed the tool's performance before go-live, how bias across demographic subgroups was evaluated, what ongoing monitoring is in place, and what happens when performance degrades. Organizations that can produce pre-deployment validation reports, bias audit results, and post-deployment monitoring data are in a stronger position than those who can only point to vendor marketing materials and FDA clearance.
Sources
- The Joint Commission and Coalition for Health AI (CHAI). Responsible Use of Artificial Intelligence in Healthcare. 2024.
- Sutter Health. AI Imaging Validation and Monitoring Program: Case Study in Multi-Site Governance. Conference presentation, 2025.
- Office of the National Coordinator for Health Information Technology (ONC). HTI-1 Final Rule: Appropriateness Criterion for AI Tools. 2024.
- Health AI Partnership (HAIP). AI Lifecycle Governance Framework with Decision Gate Methodology. 2024.
- IBM Research. AI Fairness 360: An Extensible Toolkit for Detecting, Understanding, and Mitigating Algorithmic Bias. 2019, updated 2024.
- Coalition for Health AI (CHAI). Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare. 2023.
- U.S. Food and Drug Administration. AI and ML in Medical Devices: Proposed Regulatory Framework for Modifications. Updated 2026.
- NIST AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. January 2023.
Related Questions
- ›How do we evaluate an AI vendor's claims and what questions should we ask before signing?
- ›What should we require from AI vendors as a condition of deployment?
- ›What will a Joint Commission surveyor ask about our AI governance?
- ›What is the FDA's current regulatory stance on clinical decision support software?
- ›How should we document AI-assisted decisions in the medical record?
- ›What AI governance do rural health organizations need when implementing AI through RHTP funding?

