Key Takeaways
- ·In September 2024 the Texas Attorney General announced a settlement with Pieces Technologies resolving allegations that the company made false and misleading statements about the accuracy and safety of a generative AI product used in Texas hospitals. It is a reminder that vendor accuracy claims deserve independent checking.
- ·Vendor accuracy figures often come from the vendor's own testing, on data that may not resemble your patients. Ask how the number was produced before you look at the number.
- ·The Joint Commission and CHAI guidance says organizations should ask vendors how a tool was tested and validated, and whether the vendor will tune or validate it on a sample that represents your setting. CHAI's playbook leaves it to each organization to decide, by risk tier, when local validation is needed.
- ·Four contract topics deserve a conversation with counsel before signing: support for local validation, the right to suspend use, audit and change-notification terms, and limits on how the vendor may use your data.
- ·ONC's HTI-1 rule requires developers of certified health IT to give customers information about the predictive decision support tools they supply. ONC proposed in December 2025 to remove those requirements, and the proposal hadn't been finalized when we last checked. If it is, that information will have to come through your contract.
The short answer
Start by asking for the methodology behind every accuracy claim, and look at the number last. For higher-risk tools, plan local validation on your own patient population before full deployment. Before anyone signs, ask counsel to look at suspension rights, audit and change-notification terms, and limits on data use. These are easier conversations before go-live than after a problem.
What the Texas Attorney General's Pieces Technologies Settlement Shows
The Texas Attorney General's September 2024 settlement with Pieces Technologies is worth reading for any governance committee evaluating an AI vendor. The Attorney General described it as the first settlement of its kind involving generative AI in healthcare.
According to the Attorney General's announcement, Pieces marketed the accuracy of its product by claiming a "severe hallucination rate" of "<1 per 100,000." The Attorney General's investigation found that those metrics "were likely inaccurate and may have deceived hospitals about the accuracy and safety of the company's products." The settlement resolved the allegations. Under it, Pieces agreed to accurately disclose the extent of its products' accuracy and to make sure hospital staff understand how far they should rely on the products. This summary comes from the Attorney General's announcement. It describes allegations that were settled, and no court ruled on them.
The takeaway for a governance committee is that an impressive-sounding number doesn't tell you what you need to know. Ask what was measured, on which patients, and by whom.
Patterns to Watch For in Vendor Accuracy Claims
- ·High accuracy claims measured at the data-point level, not at the clinical output level patients or clinicians see
- ·Validation studies from limited populations or single-site deployments presented as general performance
- ·Performance claims that aren't revisited when customers report errors
- ·No documentation of known failure modes or hallucination patterns shared with prospective customers
A Published Example of Why Local Checking Matters
In 2021, researchers at the University of Michigan published an external validation of the Epic Sepsis Model in JAMA Internal Medicine. The authors described it as a proprietary model "implemented at hundreds of US hospitals" whose performance "has not been adequately evaluated despite widespread use."
Across 38,455 hospitalizations at Michigan Medicine, the model's area under the curve was 0.63. It did not identify 1,709 of the 2,552 patients with sepsis (67%), while generating alerts for 18% of all hospitalized patients. The authors concluded that the model "has poor discrimination and calibration in predicting the onset of sepsis."
The study covers one model at one health system in 2018 and 2019, so it says nothing about any other product or about that model today. Its value for a governance committee is the method. An independent team measured a widely used tool on its own patients and found something different from what users had assumed.
Four Questions to Ask About Any Accuracy Claim
Governance committees approaching vendor marketing need a consistent set of questions that go beyond "what's your accuracy rate." The methodology behind the number matters more than the number itself.
How was accuracy measured?
Ask for the exact methodology. Specifically: was accuracy measured at the aggregate level, at the data-point level, or at the level of the clinical output a provider would use? A tool can have a very low error rate per individual data field while producing summaries or recommendations that are meaningfully wrong. Ask whether results were independently validated or whether the vendor conducted the validation themselves.
What populations were tested?
AI models trained and validated on one patient population may perform significantly differently on yours. If the vendor's evidence comes from a single academic medical center and you're a community hospital serving a different demographic mix, the performance figures may not translate. Ask for the sample size, the sites, and the population characteristics. If the evidence base is narrow, require in-house validation before full deployment.
What are the known failure modes and edge cases?
Any vendor with a mature safety culture maintains documentation of how their AI fails: what conditions trigger hallucinations, what patient characteristics increase error rates, what workflows produce unexpected outputs. Ask for this documentation explicitly. A vendor that claims they have none hasn't found their failure modes. A vendor that refuses to share them is telling you something about how they'd respond to a post-deployment problem. For tools built on foundation models, CHAI's Third Party Management playbook says organizations "should treat absence of this documentation as equivalent to undisclosed model limitations."
Can you provide raw validation data?
Some health systems build their own de-identified test datasets to benchmark vendor algorithms, which tells you more than a vendor-provided performance summary. Ask whether the vendor will support your team in running the same exercise. If they're confident in their product, this request shouldn't be a problem. Organizations that can't get access to validation data before signing a contract should factor that into the procurement decision.
Contract Topics to Raise With Counsel
CHAI's Third Party Management playbook says organizations "should consult qualified legal counsel when reviewing AI vendor contracts." We agree, and we aren't attorneys. These are four topics we suggest a governance committee put on counsel's list, with the operational reason each one matters.
Local validation clauses
A local validation clause commits the vendor to assist your organization in conducting an in-house performance evaluation before or during deployment, including providing access to model outputs or underlying model details sufficient to support that evaluation. Without this, you're dependent on the vendor's own validation claims.
Suspension rights
A suspension right gives your organization the ability to halt use of the AI tool without financial penalty if performance deviates materially from the vendor's stated specifications. Without it, you may be obligated to keep paying for a tool you've stopped using for safety reasons. Whether that is true of your agreement is a question for counsel.
Audit rights and performance guarantees
Audit rights allow your organization to request evidence of ongoing model performance, and change-notification terms tell you when the underlying model, its validation or its performance characteristics change. Performance terms can tie the contract to specific, measurable levels in place of the vendor's initial marketing claims. CHAI's playbook calls responsibility and timelines for incident reporting "a consequential gap in many current vendor agreements."
Data use restrictions
Under HIPAA, a vendor acting as your business associate may use patient information only as your agreement allows, and may de-identify it only to the extent the agreement authorizes. That makes the contract the place where it gets decided whether your data can be used to train or improve the vendor's models. CHAI's playbook suggests re-reading each platform's AI data processing terms, which it says "often differ materially from the general terms of service and may be updated independently of the underlying agreement," and it describes adding an AI addendum to the BAA. Ask counsel what your current agreements permit.
What Working Oversight Looks Like
Organizations that catch AI performance problems are the ones with oversight processes already in place. The practical model is to deploy with monitoring, validate continuously, and act when the evidence warrants.
Monitor after go-live
Track false alarm rates and clinician burden from the first week, and be ready to switch a tool off when the evidence says it isn't helping.
Check the vendor's numbers yourself
Compare the tool's performance on your own cases with what was marketed, and treat a large gap as a conversation to have with the vendor and with counsel.
Suspend, fix, re-validate
When post-deployment review finds missed findings, suspend use until the vendor provides a correction and your own team has validated it.
Where Federal Transparency Rules Stand
ONC's HTI-1 final rule (January 2024) requires developers of certified health IT to make information about predictive decision support tools available to their customers. Those duties fall on developers of certified health IT. They don't reach AI vendors generally, and they don't fall on hospitals. In December 2025, ONC's HTI-5 proposed rule proposed removing those requirements. It hadn't been finalized when we last checked, so the requirements are still in force. If the proposal is finalized, health systems that want that information will need to ask for it in the contract.
Frequently Asked Questions
Common questions healthcare executives ask when evaluating AI vendors.
What is the Texas Pieces Technologies settlement and why does it matter?
In September 2024 the Texas Attorney General announced a settlement with Pieces Technologies, a Dallas-based healthcare AI company, resolving allegations that the company made false and misleading statements about the accuracy and safety of a generative AI product used in several Texas hospitals. According to the Attorney General's announcement, Pieces had advertised a 'severe hallucination rate' of less than 1 per 100,000, and the investigation found those metrics were likely inaccurate and may have deceived hospitals. Pieces agreed to accurately disclose the extent of its products' accuracy and to make sure hospital staff understand how far to rely on them. These were allegations resolved by settlement, and no court ruled on them. The Attorney General called it the first settlement of its kind, and it is a useful prompt for any committee to ask how a vendor's accuracy figures were produced.
What is 'local validation' and how do we do it?
Local validation means testing a vendor's AI tool against your own patient population, using your own data, before relying on it clinically. It's distinct from accepting the vendor's published validation studies, which may have been conducted on a different population under different conditions. In practice your clinical informatics and quality teams work with the vendor to run the AI against a representative sample of your historical cases where the correct answer is already known, then compare the AI's output with the documented clinical truth. The Joint Commission and CHAI guidance suggests asking vendors during procurement whether they are willing to tune or validate on a sample that represents your setting, and CHAI's playbook leaves it to each organization to decide, by risk tier, when local validation is needed and when a vendor's attestation is enough.
Should we ever accept vendor FDA clearance as a substitute for our own evaluation?
We recommend against it. For devices cleared through the 510(k) pathway, FDA describes clearance as a finding that the device is substantially equivalent to one already legally marketed. That is a regulatory finding about the device. It doesn't measure how the tool performs on your patient population. Treat FDA status as one input to your evaluation.
What contract topics should we raise with counsel?
Four: vendor support for local validation, the right to suspend use without penalty if performance deviates from specifications, audit and notification terms covering model changes and incidents, and limits on how the vendor may use your data, including whether it may de-identify it. Those are the governance reasons to raise each topic. How the terms are drafted, and what your existing agreements already say, is work for your legal counsel.
What does 'suspension rights' mean in a vendor contract?
A suspension right is a contract term that lets your organization stop using a vendor's AI tool, and stop paying for it, without financial penalty if the tool's performance deviates materially from the vendor's stated specifications. It matters because AI model performance can degrade over time, particularly as patient populations shift or as the vendor updates the underlying model. Whether your current agreements give you that right is a question for counsel.
How do we handle AI vendors who refuse to share performance data?
Treat it as a significant factor in the procurement decision. Vendors often cite trade secrets, and there is usually room between full disclosure and nothing: a summary of validation methods and results, known limitations, and support for your own testing. If you proceed with a vendor who won't share performance information, document that decision and make sure your governance committee has explicitly accepted the risk.
Sources
- Office of the Texas Attorney General. Attorney General Ken Paxton Reaches Settlement in First-of-its-Kind Healthcare Generative AI Investigation, news release. September 18, 2024.
- U.S. Food and Drug Administration. Premarket Notification 510(k). Content current as of August 22, 2024.
- Wong A, Otles E, Donnelly JP, et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine. 2021;181(8):1065-1070.
- The Joint Commission and the Coalition for Health AI. The Responsible Use of AI in Healthcare (RUAIH), Element 4, Ongoing Quality Monitoring. September 17, 2025. (PDF)
- Coalition for Health AI. AI Governance Playbooks, Subdomain 4.4: Third Party Management, and Subdomain 4.1: Lifecycle Management. Released May 27, 2026.
- 45 CFR 164.502 (business associate uses of protected health information) and HHS Office for Civil Rights, Guidance Regarding Methods for De-identification of Protected Health Information.
- ASTP/ONC. HTI-1 Final Rule, 89 FR 1192, January 9, 2024 (45 CFR 170.315(b)(11)), and HTI-5 Proposed Rule, 90 FR 60970, December 29, 2025.

