Key Takeaways
- ·FDA clearance doesn't guarantee an AI tool performs safely on your specific patient population. A model trained on data from an academic medical center in Boston may not be appropriate for a rural clinic in Alabama.
- ·Local validation means testing vendor-reported performance against your own patient population using your own data. At HIMSS26, the chair of Sutter Health's imaging service line described tracking two objective measures, cancer detection rate and callback rate, as its breast imaging AI rolled out.
- ·Shadow deployment, running AI in silent mode alongside existing workflows before full go-live, is a widely used way to evaluate risk and bias before integration. The CHAI Lifecycle Management playbook describes silent evaluation as part of pre-deployment.
- ·We recommend stratifying bias audits by race, age, sex, and socioeconomic status, using your local patient data in addition to what the vendor reports.
- ·The Joint Commission and CHAI guidance says organizations should find out whether a tool was tested for the populations they serve and tuned or tested on local data. Separately, a federal rule, 45 CFR 92.210, calls for reasonable efforts to identify and mitigate discrimination risk from patient care decision support tools.
- ·AI upgrades require re-validation. A new model version is not the same tool you approved.
The short answer
Local validation means running the AI against your own patient population, using your own data, before relying on it clinically. It starts with a shadow deployment phase where the AI generates outputs but they're not acted upon, which lets you evaluate accuracy and bias risk without affecting patient care. It includes bias audits stratified by the demographic characteristics of your specific patient mix, using your local data in addition to what the vendor reports. It produces documented results showing the gap between vendor-claimed performance and locally validated performance. And it establishes go/no-go criteria that your governance committee decided before validation began. How much local validation a tool needs should scale with its risk, a point the CHAI Lifecycle Management playbook makes directly.
Why Vendor Validation Studies Aren't Enough
When an AI vendor presents validation evidence, they're showing you how their tool performed in the conditions they chose to study. Those conditions, population demographics, clinical environment, workflow context, and data quality, are almost never identical to yours. A tool trained and validated at a large urban academic medical center may perform meaningfully differently when deployed at a community hospital serving a different patient mix. The Joint Commission and CHAI guidance gives a plain example: a tool developed on data from mostly younger, healthy patients may perform worse for older patients.
Federal transparency rules help only partly. ONC's HTI-1 final rule (January 2024) requires developers of certified health IT to give users information about the predictive decision support tools they supply, so users can judge a tool's quality for themselves. It doesn't require providers to validate locally, and in December 2025 ONC proposed, in its HTI-5 rule, to remove those transparency requirements. FDA clearance is a separate matter. For devices that reach the market through the 510(k) pathway, clearance means FDA found the device substantially equivalent to one already legally marketed. That is a regulatory finding about the device, and it says nothing specific about how the tool performs on your patients.
The guidance treats this as a question to raise at procurement. It says organizations should ask vendors how a tool was tested and validated for its intended use, whether the vendor is willing to tune or validate on a sample that represents the deployment setting, and how relevant biases were evaluated. An organization that relied entirely on vendor-supplied evidence hasn't answered those questions for itself.
The Shadow Deployment Phase
A shadow deployment, also called silent mode deployment, runs the AI in parallel with existing workflows so that outputs are generated and visible to evaluators, but not acted upon by clinicians. This phase lets you assess AI performance against real cases without any clinical consequences, and it allows comparison against current operational standards before the tool goes live.
The shadow deployment phase serves several governance functions at once. It gives your clinical informatics team access to actual AI outputs across a representative sample of your patient population, which is the foundation for bias audits and performance analysis. It lets you identify whether the tool's output format and integration with your EHR workflows work as the vendor described. And it gives clinical staff familiarity with the tool's behavior before it influences care decisions, which improves the quality of the human review that follows go-live.
Define the evaluation period before you start
Establish how long the shadow deployment will run, what volume of cases is needed for adequate statistical power, and which patient subgroups need to be represented. Setting these parameters in advance prevents the evaluation from becoming open-ended or from stopping as soon as initial results look favorable.
Establish go/no-go criteria before you see the results
Your governance committee should define the performance thresholds that determine whether the tool proceeds to full deployment before validation begins. Defining criteria after you've already seen promising results, or after you've invested in a vendor relationship, creates pressure to rationalize rather than evaluate.
Use the shadow phase to assess workflow integration
Shadow deployment surfaces integration problems that weren't visible in vendor demos: whether outputs display correctly in your EHR, whether alert fatigue is a concern, whether the workflow assumptions built into the tool match how your clinicians work.
Document the comparison against current standards
Evaluate shadow deployment outputs against current clinical performance on the same cases where you have historical data. The question is whether the AI performs better than what you're already doing, and by how much, and whether that improvement holds across all patient subgroups or only in aggregate.
Bias Audits: What They Need to Cover
AI tools can perform well in aggregate while performing materially worse for specific patient subgroups. An imaging AI with strong overall sensitivity might have significantly lower sensitivity for patients of a particular age or body type. A sepsis prediction model might generate higher false positive rates for patients from certain demographic backgrounds. Aggregate accuracy metrics hide these disparities.
The Joint Commission and CHAI guidance (RUAIH) addresses this in Element 6, Risk and Bias Assessment. It says organizations should determine whether algorithms "are tested for the specific populations they serve and ensure they are appropriately tuned and/or tested on local data." It adds that, in addition to vendor-reported information, organizations should check for bias when validating on local data and again after deployment, either internally or with their vendors. The guidance is voluntary and doesn't list demographic categories. The dimensions below are our recommendation.
One binding federal rule sits next to that guidance. Under 45 CFR 92.210, a covered entity "must not discriminate on the basis of race, color, national origin, sex, age, or disability in its health programs or activities through the use of patient care decision support tools." The same rule sets "an ongoing duty to make reasonable efforts to identify" tools that use those characteristics as inputs, and to make reasonable efforts to mitigate the resulting risk. Those two duties applied from May 1, 2025. Ask your attorney how the rule applies to your organization and whether anything about its enforcement has changed.
Practical tools for running these analyses include AI Fairness 360, an open-source toolkit first released by IBM Research, and Google's Fairness Indicators. Both are designed to measure whether AI performance differs systematically across demographic subgroups. The analysis requires access to outcome data that lets you compare AI outputs to known clinical results, which is why bias audits work best when conducted on historical cases where the correct answer is already documented.
Bias Audit Dimensions to Evaluate
- ·Race and ethnicity: performance parity across all racial groups represented in your patient population
- ·Age: accuracy at the extremes of age distribution, particularly pediatric and elderly populations
- ·Sex and gender: whether performance differs between male and female patients, and how the tool handles non-binary gender data
- ·Socioeconomic status: whether outcomes proxy variables in the training data introduced bias against lower-income patients
- ·Comorbidity patterns: whether the tool performs differently for patients with multiple conditions versus primary diagnoses
- ·Geographic and facility type: if validating across multiple sites, whether performance is consistent or varies by location
A Reference Point: Sutter Health's Objective Measures
At HIMSS26 in March 2026, the chair of Sutter Health's imaging service line described how the system evaluated its breast imaging AI. Sutter didn't lean on clinicians' impressions of the tool. It tracked two objective measures, cancer detection rate and callback rate, through a pilot and then a wider rollout, and kept tracking them after go-live.
When the tool moved from one major version to the next, Sutter tested the new version before relying on it. The principle applies broadly. A new model version can have different training data, different performance, and different bias patterns, so approval of one version shouldn't carry over to the next automatically. The Joint Commission and CHAI guidance makes a similar observation: AI tools and their underlying algorithms "may be updated periodically," which means performance can change.
Ongoing measurement also addresses the performance drift problem. AI tools can degrade over time as patient populations shift, as data quality changes, or as clinical workflows evolve in ways that alter the distribution of inputs the AI receives. Continuous monitoring with objective metrics provides early warning of drift before it becomes a patient safety issue.
Pre-deployment
- ·Shadow deployment with objective KPIs
- ·Bias audit stratified by local demographics
- ·Go/no-go criteria set in advance
- ·Gap analysis vs. vendor-claimed performance
Go-live governance
- ·Defined human review standards
- ·Escalation process for AI errors
- ·Integration monitoring for workflow issues
- ·Initial performance reporting to governance committee
Ongoing monitoring
- ·Monthly or quarterly KPI review
- ·Drift detection against baseline metrics
- ·Re-validation triggered by model updates
- ·Periodic bias re-audits as population shifts
What to Document
This is the record your governance committee and board should be able to produce:
Pre-deployment validation report
The performance metrics from your shadow deployment phase, including overall accuracy against your patient population, results stratified by the demographic dimensions of your bias audit, the gap between vendor-claimed performance and locally measured performance, and the governance committee's go/no-go decision with rationale.
Bias audit results
Documentation of the demographic subgroups evaluated, the metrics used to assess disparate impact, the specific results for each subgroup, and any remediation decisions made when disparities were identified. The audits should be conducted on your local population data, and the documentation should make that explicit.
Go-live criteria and approval record
The performance thresholds your governance committee established before validation began, evidence that validation results met those thresholds, and the formal governance committee approval authorizing deployment. This record should predate go-live.
Ongoing monitoring reports
Post-deployment performance data showing how the tool is performing against the KPIs established at go-live, including trend data showing whether performance is stable or drifting, and any exceptions or incidents that triggered governance review.
Re-validation records for model updates
For each material vendor update to the AI model, documentation of the re-validation process, results, and governance committee approval before the update was deployed in your environment.
The Health AI Partnership's Decision Points
One of the Health AI Partnership's key decision points for adopting an AI tool is "Generate evidence of safety, efficacy and equity." Treat that as a real decision made before go-live, so deployment doesn't proceed by default because procurement has gone too far to reconsider. The governance committee decides on the evidence in front of it.
Frequently Asked Questions
Common questions from health system leaders building pre-deployment validation processes.
What is local validation and how is it different from vendor validation?
Local validation means testing an AI tool's performance against your own patient population, using your own data, before relying on it clinically. Vendor validation is testing the vendor conducted on the population they chose, under conditions they controlled. The two can differ. A tool validated at a large urban academic medical center may perform quite differently at a community hospital serving a different demographic mix. Local validation closes that gap by confirming the tool works for your patients.
Does FDA clearance mean we don't need to validate locally?
We recommend validating locally either way. For devices cleared through the 510(k) pathway, FDA describes clearance as a finding that the device is substantially equivalent to one already legally marketed. That is a regulatory finding about the device. It doesn't measure how the tool performs on your patient population. Treat FDA status as one input to your governance review.
What is a shadow deployment and how do we run one?
A shadow deployment, also called silent mode, runs an AI tool in parallel with existing workflows. The AI generates outputs that are visible to your evaluation team, but those outputs don't influence clinical decisions during the shadow phase. This lets you assess how the tool performs on real cases from your patient population without clinical consequences if the AI makes errors. Shadow deployment should run long enough to capture a representative sample of your patient mix and use cases, with performance evaluated against historical cases where the correct clinical answer is already documented.
What does a bias audit for an AI tool involve?
A bias audit assesses whether an AI tool performs differently across demographic subgroups. The audit compares AI performance on cases from different patient populations and asks whether accuracy, sensitivity, specificity, or false positive rates differ systematically by race, ethnicity, age, sex, or socioeconomic characteristics. We recommend running it on your local patient data, in addition to reviewing what the vendor reports, because demographic patterns in your patient mix may differ from those in the vendor's validation dataset. Open-source tools such as AI Fairness 360 support this analysis.
Do we need to re-validate when a vendor updates their AI model?
We recommend it. A new version can have different training data, different performance, and different bias patterns, so the approval that covered one version shouldn't automatically extend to the next. At HIMSS26, Sutter Health described testing a new major version of its breast imaging AI before relying on it. This is also why contracts need model change notification terms: you can't re-validate an update nobody told you about.
What do the Joint Commission and CHAI say about validating AI tools?
Their joint guidance (RUAIH, September 2025) is voluntary. It says organizations should ask vendors how a tool was tested and validated, find out whether it was tested for the populations they serve and tuned or tested on local data, check for bias before deployment and afterward, and keep monitoring performance on an ongoing, risk-based basis. On June 1, 2026 the Joint Commission began offering a voluntary certification based on that guidance, and accreditation isn't needed to apply. Organizations that can produce pre-deployment validation reports, bias audit results, and monitoring data are in a stronger position than those that can point only to vendor materials and FDA clearance.
Sources
- The Joint Commission and the Coalition for Health AI. The Responsible Use of AI in Healthcare (RUAIH), Elements 4 and 6. September 17, 2025. (PDF)
- The Joint Commission. Voluntary Responsible Use of AI in Healthcare Certification, news release. June 1, 2026.
- Coalition for Health AI. AI Governance Playbooks, Subdomain 4.1: Lifecycle Management. Released May 27, 2026.
- 45 CFR 92.210, Nondiscrimination in the use of patient care decision support tools, and 45 CFR 92.1 (applicability dates). eCFR, accessed September 19, 2026.
- ASTP/ONC. HTI-1 Final Rule, 89 FR 1192, January 9, 2024 (45 CFR 170.315(b)(11)), and HTI-5 Proposed Rule, 90 FR 60970, December 29, 2025.
- U.S. Food and Drug Administration. Premarket Notification 510(k). Content current as of August 22, 2024.
- Health AI Partnership. Key Decisions in Adopting an AI Solution. Accessed September 19, 2026.
- Bellamy RKE, et al. AI Fairness 360: An Extensible Toolkit for Detecting, Understanding, and Mitigating Unwanted Algorithmic Bias. arXiv:1810.01943, October 2018.
- Kiesner J. Presentation on imaging AI for cancer detection at Sutter Health, HIMSS26, Las Vegas, March 11, 2026. Authors' notes from attendance.
Related Questions
- ›How do we evaluate an AI vendor's claims and what questions should we ask before signing?
- ›What should we require from AI vendors as a condition of deployment?
- ›What will a Joint Commission surveyor ask about our AI governance?
- ›What is the FDA's current regulatory stance on clinical decision support software?
- ›How should we document AI-assisted decisions in the medical record?
- ›What AI governance do rural health organizations need when implementing AI through RHTP funding?

