Healthcare AI validation gap: 93% deploy, 44% can test
New UPMC and KLAS data show health systems have mainstreamed third-party AI faster than they built places to test it, leaving CIOs to own a safety burden regulators are shifting downstream.

The healthcare AI validation gap is now measurable, and it is wide. New data from the Center for Connected Medicine at UPMC and KLAS Research, released Aug. 6, 2026, found that 93% of surveyed health system leaders say their organization deploys third-party AI, while only 44% have a dedicated data platform or environment for testing those tools. Adoption has outrun oversight, and the accountability for what happens in between is landing on health IT leadership.
Black Book Research survey of 182 hospital leaders, Nov. 11, 2025, via Becker's.
| Value (% of IT and quality/safety budget) | Median budget share |
|---|---|
| All hospitals (median) | 4.2 % of IT and quality/safety budget |
| Large health systems | 6.8 % of IT and quality/safety budget |
| Small hospitals | 2.3 % of IT and quality/safety budget |
What the UPMC and KLAS numbers actually say
The survey, which spanned executives at 27 health systems, paints a picture of near-universal deployment paired with uneven infrastructure. Clinical documentation was the most commonly cited AI deployment area at 52%, which means ambient documentation is often the highest-volume AI surface in a health system and, in many cases, the least formally validated. Roughly 63% of organizations described their AI governance as "developing" or "ad hoc."
That combination matters because documentation AI touches nearly every encounter. A model that summarizes a visit inaccurately does not trigger an alarm the way a failed interface does. It quietly enters the record, the claim, and eventually the litigation file. Without a test environment, the first real evaluation of the tool is production.
Lighter vendor certification does not eliminate validation work. It relocates it - to the health system.
Why HTI-5 shifts the validation burden onto providers
HHS ASTP/ONC's HTI-5 proposed rule, announced Dec. 22, 2025 and published Dec. 29, 2025, simplifies certification requirements for AI models and pushes FHIR-first exchange while tightening information blocking exceptions. Legal and policy analyses from ReedSmith and McDermott, along with coverage in Becker's and Healthcare IT News, read the direction consistently: vendors face a lighter certification path.
Lighter vendor certification does not eliminate the validation work. It relocates it. If a model arrives with fewer federally mandated attestations, the provider organization becomes the last party positioned to determine whether it performs on its own population, in its own workflows, with its own data quality problems. That is not a compliance footnote. It is the new patient safety line, and boards will eventually ask who signed off.
Shadow AI and the pressure that creates it
The governance gap is not only about procured tools. A Wolters Kluwer Health survey of more than 500 professionals, published Jan. 22, 2026, found that 40% of healthcare professionals have encountered unauthorized "shadow AI" tools at work and roughly 17% to 20% admit using them. More than half of administrators cited faster workflows as the reason. Staff are not defying policy for sport. They are routing around slow processes.
The demand pressure on leadership is equally real. Qventus' April 2026 "Beyond the Pilot" report found that 94% of healthcare CIOs believe AI delays create competitive disadvantage. That figure explains a great deal of the current behavior: pilots move past governance because the perceived cost of waiting is higher than the perceived cost of an unvalidated deployment. Until that calculation changes, the validation gap will persist.
What health systems are spending on AI safety
Budget allocation reveals priorities more honestly than strategy decks. Black Book Research's survey of 182 hospital leaders, published Nov. 11, 2025, found the median share of 2026 IT and quality and safety budgets dedicated to AI governance and safety was 4.2%. Large health systems reported 6.8%. Small hospitals reported 2.3%.
The spread is the story. Large systems can build validation infrastructure and staff a chief AI officer function. Small and rural hospitals deploying the same third-party tools are working with roughly a third of that proportional investment, and they buy from many of the same vendors. If HTI-5 lightens vendor certification, the organizations least able to validate locally carry the most residual risk.
Building a credible validation stack, and naming its owner
UPMC's own answer is instructive. The system built Ahavi, a real-world data platform used to validate third-party models "in silico" on de-identified data before deployment. The principle is transferable even where the budget is not: test the model against your population before it reaches a patient, and keep the evidence.
A defensible stack has four parts. First, a complete model inventory that includes AI embedded in EHR modules and point solutions, not just standalone purchases. Second, a pre-deployment test environment with representative de-identified data, or a documented shared arrangement if building one is out of reach. Third, post-deployment monitoring with defined performance thresholds and a rollback path. Fourth, a named accountable owner, whether that is the CMIO, a chief AI officer, or a standing governance committee with real authority to say no.
The final piece is the least technical and most often skipped. Ad hoc governance usually means no one person can answer who approved a given model, on what evidence, and when it was last reviewed. That question will arrive from a board, a payer disputing a documentation-driven claim, or opposing counsel. Having an answer ready is now part of the job.


