Key Takeaways
- ·The Texas AG settlement with Pieces Technologies is the enforcement benchmark: vendors who market AI with misleading accuracy claims face regulatory action, even when hospitals signed off on the deployment.
- ·Vendor accuracy metrics are almost always measured under ideal conditions. The methodology matters more than the number.
- ·Local validation — testing vendor performance against your own patient population with your own data — is now an organizational responsibility, not optional.
- ·Most AI vendor contracts missing critical protections: suspension rights, audit rights, notification requirements for model changes, and data use restrictions.
- ·HTI-5's proposed removal of federal model card requirements shifts evaluation responsibility to health systems. Federal deregulation doesn't reduce your accountability — it increases it.
The short answer
Start by demanding the methodology behind every accuracy claim — not just the number. Then require local validation on your own patient population before full deployment. And before you sign anything, check the contract for suspension rights, audit rights, and data use restrictions. Most contracts that arrive from AI vendors are missing at least two of those three. The organizations that discover this after an adverse event are in a far worse position than those who negotiate before go-live.
The Enforcement Benchmark: What Happened with Pieces Technologies
The Texas Attorney General's settlement with Pieces Technologies is required reading for any health system governance committee evaluating an AI vendor. It establishes what misleading AI marketing looks like in a healthcare context, and it signals that regulators are watching.
Pieces Technologies marketed its AI clinical documentation tool with claims including an error rate below 1 in 100,000 and assertions that it performed "better than humans in a lot of cases." Both claims were problematic. The accuracy methodology measured performance at the data-point level — not at the level of the clinical summary a physician would actually rely on. The AI produced hallucinations, meaning it fabricated information that wasn't in the source record. And the company continued marketing those same claims after receiving reports of errors from hospital clients.
The lesson isn't that AI documentation tools are inherently dangerous. It's that impressive-sounding numbers don't tell you what you need to know, and governance committees that take vendor metrics at face value are operating without adequate protection.
The Pieces Technologies Pattern to Watch For
- ·High accuracy claims measured at the data-point level, not at the clinical output level patients or clinicians see
- ·Validation studies from limited populations or single-site deployments presented as general performance
- ·Continued marketing of performance claims after receiving error reports
- ·No documentation of known failure modes or hallucination patterns shared with prospective customers
Four Questions to Ask About Any Accuracy Claim
Governance committees approaching vendor marketing need a consistent set of questions that go beyond "what's your accuracy rate." The methodology behind the number matters more than the number itself.
How was accuracy measured?
Demand the exact methodology. Specifically: was accuracy measured at the aggregate level, at the data-point level, or at the level of the clinical output a provider would actually use? A tool can have a very low error rate per individual data field while producing summaries or recommendations that are meaningfully wrong. Ask whether results were independently validated or whether the vendor conducted the validation themselves.
What populations were tested?
AI models trained and validated on one patient population may perform significantly differently on yours. If the vendor's evidence comes from a single academic medical center and you're a community hospital serving a different demographic mix, the performance figures may not translate. Ask for the sample size, the sites, and the population characteristics. If the evidence base is narrow, require in-house validation before full deployment.
What are the known failure modes and edge cases?
Any vendor with a mature safety culture maintains documentation of how their AI fails — what conditions trigger hallucinations, what patient characteristics increase error rates, what workflows produce unexpected outputs. Ask for this documentation explicitly. A vendor that claims they have none hasn't found their failure modes. A vendor that refuses to share them is telling you something important about how they'd respond to a post-deployment problem.
Can you provide raw validation data?
Emory Healthcare built their own de-identified test datasets to benchmark vendor algorithms rather than accepting vendor-provided performance summaries. That's the gold standard. Ask whether the vendor will support your team in running the same exercise. If they're confident in their product, this request shouldn't be a problem. Organizations that can't get access to validation data before signing a contract should factor that into the procurement decision.
Contract Protections That Are Often Missing
Most AI vendor contracts that arrive from the vendor's legal team are written to protect the vendor. The provisions that protect your organization — and that most commonly aren't there — fall into four categories.
Local validation clauses
A local validation clause commits the vendor to assist your organization in conducting an in-house performance evaluation before or during deployment, including providing access to model outputs or underlying model details sufficient to support that evaluation. Without this, you're dependent on the vendor's own validation claims.
Suspension rights
A suspension right gives your organization the ability to halt use of the AI tool without financial penalty if performance deviates materially from the vendor's stated specifications. This is the provision that UPMC exercised when it discontinued an early AI tool due to excessive false alarms, and that Atrius Health used when validation showed a tool performing only marginally better than chance. Without it, you may be contractually obligated to continue paying for a tool you've stopped using for safety reasons.
Audit rights and performance guarantees
Audit rights allow your organization to request evidence of ongoing model performance, including notification of any changes to the underlying model, validation methodology, or performance characteristics. Performance guarantees tie contract terms to specific, measurable performance levels — not just the vendor's initial marketing claims. These should include the right to renegotiate or exit if performance falls below the contractually specified threshold.
Data use restrictions
AI vendors frequently include provisions allowing them to use de-identified patient data to improve their models. This needs explicit review and, in most cases, explicit restriction. Patient data should not be used for vendor model improvement without patient consent and appropriate legal authorization. This requires a current Business Associate Agreement that addresses AI-specific data use, not just a standard BAA written before AI training data was a consideration.
Real Examples of Governance Working
The organizations that discovered AI performance problems were able to act because they had oversight processes in place. The pattern — deploy with monitoring, validate continuously, act when evidence warrants — is the practical model.
UPMC
Discontinued an early AI tool after post-deployment monitoring identified excessive false alarm rates that were creating clinician burden rather than reducing it.
Atrius Health
Cancelled a vendor contract when their own validation showed the tool performed only marginally better than chance — well below the vendor's marketed performance claims.
Baptist Health
Halted a vascular ultrasound AI that was missing aneurysms in post-deployment review. Use was suspended until the vendor provided an update and Baptist Health's team validated the correction.
The Federal Deregulation Shift
ONC's proposed HTI-5 rule would have removed federal model card requirements for AI used in certified health IT. That shift transfers evaluation responsibility to health systems. Federal deregulation doesn't reduce organizational accountability for AI safety — it increases it, by removing the federal floor that once provided some assurance about what vendors were required to disclose. Organizations that relied on regulatory requirements to ensure vendor transparency will need internal processes to fill that gap.
Frequently Asked Questions
Common questions healthcare executives ask when evaluating AI vendors.
What is the Texas Pieces Technologies settlement and why does it matter?
The Texas Attorney General's office settled with Pieces Technologies, a healthcare AI documentation company, over allegations that the company marketed its AI with materially misleading accuracy claims. Specifically, the company claimed an error rate below 1 in 100,000 and that its AI performed better than humans — claims that were based on a methodology measuring individual data-point accuracy rather than the clinical summary accuracy that actually matters clinically. The AI also produced hallucinations, and the company continued using those marketing claims after receiving error reports from clients. The settlement matters because it's the first major enforcement action establishing what misleading AI marketing in healthcare looks like, and it signals that regulators are willing to act. Every governance committee evaluating a vendor should read it.
What is 'local validation' and how do we do it?
Local validation means testing a vendor's AI tool against your own patient population, using your own data, before relying on it clinically. It's distinct from accepting the vendor's published validation studies, which may have been conducted on a different population under different conditions. Emory Healthcare is the frequently cited example — they built de-identified test datasets from their own patient records and used them to benchmark vendor algorithms rather than accepting vendor-provided summaries. The practical implementation involves your clinical informatics and quality teams working with the vendor to run the AI against a representative sample of your historical cases where the correct answer is already known, then comparing the AI's output to the documented clinical truth. This requires vendor cooperation, which is why a local validation clause in the contract matters.
Should we ever accept vendor FDA clearance as a substitute for our own evaluation?
No. FDA clearance for a medical device means the FDA determined the device is substantially equivalent to a predicate device — it's a regulatory threshold, not a clinical performance guarantee for your specific patient population. FDA-cleared AI tools have been recalled, suspended, and discontinued by hospitals after post-deployment validation revealed performance problems. FDA clearance is one input to your evaluation, not a substitute for it. The Baptist Health vascular ultrasound case is instructive here — an FDA-cleared tool was missing aneurysms in Baptist Health's patient population. Regulatory clearance didn't prevent the problem, and Baptist Health's own monitoring detected it.
What contract protections are most commonly missing from AI vendor agreements?
Four consistently: local validation clauses (vendor commitment to support in-house performance evaluation), suspension rights (ability to halt use without penalty if performance deviates from specifications), notification requirements (vendor obligation to inform the organization of model changes, issues, or material performance changes), and data use restrictions (explicit prohibition on using patient data for vendor model improvement without consent, backed by a current and AI-specific Business Associate Agreement). Most standard vendor contracts are written to protect the vendor. The burden is on your legal counsel and governance committee to negotiate these provisions before signing.
What does 'suspension rights' mean in a vendor contract?
A suspension right is a contractual provision that allows your organization to stop using a vendor's AI tool — and stop paying for it — without financial penalty if the tool's performance deviates materially from the vendor's stated specifications. Without this provision, you may be contractually obligated to continue paying for a contract even after you've stopped using the tool for patient safety reasons. This matters because AI model performance can degrade over time, particularly as patient populations shift or as the underlying model is updated by the vendor. Having the right to suspend use if performance falls below a contractually defined threshold is a basic protection that many organizations don't have.
How do we handle AI vendors who refuse to share performance data?
Refusal to share validation data or failure mode documentation should be a significant red flag in your procurement evaluation. A vendor with genuine confidence in their product's safety and performance should welcome local validation — it protects them from liability and builds customer confidence. When vendors decline to share performance data, the most likely explanations are that the data doesn't support their marketing claims, or that the vendor's legal team has assessed the disclosure risk and decided the risk of sharing outweighs the sales benefit. Neither is a comfortable position for a health system to be in. If you're proceeding with a vendor who won't share performance data, document that decision carefully and ensure your governance committee has explicitly accepted that risk.
Sources
- State of Texas v. Pieces Technologies, Inc. Office of the Texas Attorney General. 2024.
- Emory Healthcare. "AI Vendor Evaluation and Local Validation Protocol." Internal case study cited in CHIME AI Governance guidance. 2024.
- UPMC Center for Connected Medicine. "AI Governance in Health Systems: Lessons from Early Deployments." 2023.
- Office of the National Coordinator for Health Information Technology (ONC). "HTI-5 Proposed Rule." 2025.
- American Hospital Association. "Trustworthy AI in Health Care: A Framework for AI Governance." 2024.
- Coalition for Health AI (CHAI). "Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare." 2023.
- NIST AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. January 2023.

