Key Takeaways
- Most U.S. hospitals now employ third‑party AI tools, yet fewer than 50 % maintain a dedicated environment for pre‑deployment testing.
- Validation practices vary widely—from formal vendor trials to informal pilots—producing inconsistent results.
- About 63 % of health systems describe their AI strategy as still developing or ad hoc, hampered by limited time, capital, and talent.
- Without reliable testing, organizations often embark on six‑month‑plus implementation cycles that fail to deliver expected value.
- UPMC’s Ahavi platform offers real‑world data validation against de‑identified patient records, but ongoing governance is equally critical.
- Continuous post‑implementation monitoring helps detect bias, model drift, and ensures clinical algorithms remain accurate over time.
- Experts urge health systems to embed equity evaluations into AI assessment, training clinicians to flag potential disparities.
- Clinician use of AI tools nearly doubled from 2023 to 2026, outpacing the development of standardized evaluation frameworks and leaving each organization to self‑govern AI largely on its own.
The Surge in AI Adoption Across Health Systems
Health systems nationwide have embraced artificial intelligence with unprecedented enthusiasm, deploying third‑party tools for everything from radiology triage to supply‑chain optimization. The momentum reflects a belief that AI can alleviate clinician burnout, improve diagnostic accuracy, and reduce operational costs. However, this rapid uptake has outpaced the governance structures and technical infrastructure needed to manage AI safely and effectively. As one industry observer noted, “Health systems are more committed than ever to AI deployment, with most hospitals across the country now using third‑party tools for various clinical and administrative workflows. But new research shows that this enthusiasm has outpaced the governance and infrastructure needed to effectively manage this AI surge.”
Testing Infrastructure Falls Short
A joint report from UPMC’s Center for Connected Medicine and KLAS Research reveals a stark gap: less than half of hospitals possess a dedicated environment for testing AI tools before those tools reach patient care. Without a sandbox or controlled setting, organizations risk introducing unvetted algorithms directly into clinical workflows. The report’s finding is sobering: “less than half of hospitals have a dedicated environment for testing AI tools before they touch patient care.” This deficiency leaves health systems vulnerable to performance failures, unintended biases, and regulatory scrutiny.
Variable Validation Approaches
While most organizations do evaluate AI solutions prior to rollout, the methods they employ are far from uniform. Some health systems conduct rigorous, vendor‑led validation studies; others rely on loosely structured pilot programs or ad‑hoc clinician feedback. This heterogeneity makes it difficult to compare outcomes across institutions and to establish best‑practice benchmarks. As the report observes, “While most organizations evaluate AI solutions before they’re rolled out, those validation methods vary widely, from formal vendor testing to loosely structured pilot programs.”
Strategic Immaturity and Resource Constraints
Compounding the testing shortfall, 63 % of health systems characterize their AI strategy as still developing or ad hoc. Leaders cite pressing constraints—limited budgets, scarce data‑science talent, and competing priorities—as reasons they cannot invest in robust testing pipelines up front. Ken Howard, vice president of technology services engineering at UPMC Enterprises, explains the pragmatic reality: “When a hospital identifies a challenge it believes AI can solve, it often doesn’t have the time, capital or talent needed to build a structured testing environment first — so the organization defaults to a standard IT implementation process instead.”
Consequences of Inadequate Testing
The fallback to conventional IT rollout typically stretches implementation timelines to six months or more. Yet, because the underlying AI model has not been stress‑tested in the hospital’s own data ecosystem, many projects fail to deliver the anticipated value. Howard warns of the costly cycle that results: “Without having a dedicated or consistent test environment strategy, they’re going to go down that path just to learn that all that work potentially wasn’t justified.” The mismatch between expectation and outcome erodes trust in AI and can deter future innovation.
UPMC’s Ahavi Platform as a Partial Solution
To address the testing void, UPMC has developed Ahavi, a real‑world data platform that enables hospitals to validate third‑party AI algorithms against de‑identified patient data before formal deployment. By simulating real‑clinical conditions, Ahavi helps organizations gauge performance, safety, and fit within their specific workflows. Rob Bart, UPMC’s chief medical information officer, highlights its role: “UPMC has tried to help solve this problem with Ahavi, its real‑world data platform that lets the hospitals validate third-party AI tools against de-identified patient data before they’re formally deployed.” He cautions, however, that pre‑deployment testing is only one piece of the puzzle.
Beyond Pre‑Deployment: Ongoing Governance
Bart stresses that effective AI management must extend well beyond the initial launch. UPMC has maintained a formal AI governance structure for more than two years, which includes continuous monitoring of tools after they go live. This ongoing oversight ensures that any drift in performance or emerging safety concerns are caught early. As Bart put it, “He noted that governance must extend well beyond a solution’s initial rollout — adding that UPMC has had a formal AI governance structure in place for more than two years, which involves the continuous monitoring of tools after they go live.”
Importance of Monitoring Clinical Algorithms
The need for vigilant post‑implementation scrutiny is especially acute for clinical algorithms that influence patient decisions, such as models predicting hospital length of stay or readmission risk. Bart explains the routine checks: “We monitor that on regular intervals post-implementation to make sure that the guidance that it is intended to provide is still accurate and reflective of the original [validation].” Regular recalibration against updated clinical data helps preserve the reliability of these decision‑support tools over time.
Population‑Specific Testing to Catch Bias and Drift
In addition to temporal monitoring, UPMC subjects vendor algorithms to testing against its own patient population rather than relying solely on the vendor’s generic validation datasets. This approach surfaces issues like bias, model drift, or population‑specific performance gaps that might remain hidden in broader studies. Bart adds, “UPMC also tests vendor algorithms against its own patient population, rather than relying solely on a vendor’s testing data, Bart added. This helps catch issues, including bias and model drift, that a more generic data set might miss.”
Clinician AI Use Accelerates, Standards Lag
The pace at which clinicians are adopting AI further intensifies the pressure on health systems to establish robust evaluation processes. Kate Eisenberg, senior medical director of DynaMed, cites an American Medical Association survey showing that clinician use of AI tools nearly doubled from 2023 to 2026. “Eisenberg pointed to a survey from the American Medical Association that showed clinicians’ use of AI tools nearly doubled from 2023 to 2026 — a pace of change that makes it that much harder for evaluation standards to keep up.” The rapid adoption leaves evaluation frameworks scrambling to catch up, increasing the risk of inconsistent or insufficient oversight.
Embedding Equity Evaluation into AI Assessment
Eisenberg urges health systems to go beyond technical performance and explicitly examine equity impacts when assessing AI. She notes that many platforms already provide mechanisms for users to flag equity concerns, but clinicians need targeted training to leverage them effectively. “We’ve always had in our web interface the opportunity for users to flag if there was an equity concern or not,” she said, noting that clinical teams should be specifically trained to evaluate AI responses for bias. Incorporating equity checks into routine evaluation helps prevent algorithms from exacerbating existing disparities.
The Path Forward: Self‑Governance Until Standards Emerge
Until industry‑wide standards for AI validation and governance mature, health systems must largely rely on internal policies to manage AI responsibly. The combination of pre‑deployment testing platforms like Ahavi, rigorous post‑implementation monitoring, and equity‑focused clinician training offers a pragmatic pathway. As the sector continues to innovate, aligning enthusiasm with disciplined oversight will be essential to ensure that AI delivers on its promise of safer, more equitable, and more efficient care.
Hospitals Are All In on AI, but Testing and Oversight Haven’t Caught Up

