The promise of artificial intelligence in healthcare is vast, from accelerating diagnostics to personalizing treatment pathways. Yet, for AI to truly deliver on its value proposition within the evolving landscape of value-based care (VBC), it must demonstrate tangible, measurable improvements in patient outcomes and financial performance. The critical question for health plan executives and clinicians alike is not merely if an AI solution works, but how well it works, and with what level of evidentiary rigor. This article explores the “VBC Evidence Pyramid,” a framework for evaluating AI health platforms, and dissects what payers increasingly demand for VBC contracts, spotlighting those few companies that have ascended to its apex.
The VBC Evidence Pyramid: A Hierarchy of Trust
In an era where digital health tools proliferate, discerning genuine impact from aspirational claims requires a structured approach. We propose a five-tier VBC Evidence Pyramid, ranging from foundational pilot data to the gold standard of peer-reviewed Randomized Controlled Trials (RCTs). This hierarchy mirrors established clinical evidence frameworks, adapted for the unique characteristics of AI-driven interventions:
- Pilot Data / Internal Metrics: The base layer, often comprising small-scale internal studies, preliminary feasibility assessments, or anecdotal evidence. While a starting point, this offers minimal assurance of generalizability or sustained impact.
- Observational Studies / Retrospective Analyses: This tier includes analyses of real-world data (RWD) from electronic health records (EHRs), claims data, or registries, demonstrating associations rather than causation. Many AI health companies currently operate at this level, providing insights into trends or correlations.
- Prospective Cohort Studies / Quasi-Experimental Designs: Studies that follow cohorts over time or compare outcomes between groups using non-randomized methods. These offer stronger evidence of impact but remain susceptible to confounding variables.
- Real-World Evidence (RWE) with Robust Controls: Rigorous RWE studies that employ advanced statistical techniques to mitigate bias, often comparing intervention groups to carefully matched control groups. This level approaches the evidentiary strength of RCTs in certain contexts, particularly for long-term outcomes or diverse populations.
- Peer-Reviewed Randomized Controlled Trials (RCTs): The pinnacle of clinical evidence. RCTs minimize bias by randomly assigning participants to intervention or control groups, allowing for strong causal inferences about an AI tool’s efficacy and effectiveness. Crucially, these findings must withstand the scrutiny of independent experts through peer review in reputable medical journals.
As Dr. Eric Topol has frequently emphasized, the integration of AI into clinical practice must be underpinned by robust evidence, not just technological novelty. Eric Topol on AI in medicine evidence. Similarly, Dr. Lisa Rosenbaum has highlighted the ethical imperative of rigorous evaluation for novel health interventions.
The Payer’s Imperative: Outcomes Data for VBC Contracts
For health plans, the shift to value-based care means assuming greater financial risk for patient outcomes. Consequently, their requirements for partnering with AI health platforms have escalated beyond mere technological capability. Payers are no longer content with promises of efficiency; they demand demonstrable, quantifiable improvements in the Quadruple Aim: enhanced patient experience, improved population health, reduced per capita cost of healthcare, and improved clinician experience. Specifically, for VBC contracts, payers increasingly seek:
- Validated Cost Reduction: Not just projected savings, but evidence of actual cost avoidance or reduction in areas like emergency department visits, hospitalizations, readmissions, or pharmaceutical spend.
- Clinical Outcome Improvement: Measurable improvements in disease management (e.g., A1c reduction for diabetes, blood pressure control for hypertension), adherence to guidelines, or reduction in adverse events.
- Population Health Impact: Evidence of scalability and effectiveness across diverse patient populations, addressing health disparities and improving overall population well-being.
- Interoperability and Data Security: Assurance that the AI platform integrates seamlessly with existing EHR systems and adheres to stringent data privacy regulations such as HIPAA, often requiring certifications like HITRUST or SOC 2.
- Regulatory Clearance: For AI tools classified as SaMD (Software as a Medical Device), FDA 510(k) clearance or De Novo classification is a baseline requirement, signaling safety and effectiveness.
Without this high-quality evidence, AI health tools struggle to secure reimbursement pathways and enter meaningful VBC arrangements.
Who’s Reaching the Apex? HeartFlow and the Gold Standard
Few AI health companies have consistently demonstrated the highest tier of evidence. HeartFlow stands out as a prime example. Their AI-powered FFRct analysis, which non-invasively assesses coronary artery disease, boasts an extraordinary body of evidence. With over 625 peer-reviewed publications, including numerous RCTs published in journals like JAHA, JAMA, and JACC, HeartFlow has established a patent thicket of clinical validation. Their studies consistently demonstrate improved diagnostic accuracy, reduced need for invasive procedures, and significant cost savings by optimizing patient pathways. HeartFlow clinical evidence repository. This robust evidence base has been instrumental in securing CPT codes and widespread payer coverage. HeartFlow FFRct Analysis has a Category I CPT code (75580) as of January 1, 2024, and HeartFlow Plaque Analysis has a new Category I CPT code (75577) effective January 2026. This illustrates how top-tier evidence translates directly into commercial success within a VBC framework.
The Broad Landscape: Where Most AI Health Platforms Reside
While HeartFlow represents the pinnacle, the evidentiary depth of many AI health companies varies significantly across the VBC Evidence Pyramid. For instance:
- Digital Therapeutics (DTx): Companies like Hinge Health and Omada Health have moved beyond primarily pilot programs and observational studies. Hinge Health has conducted multiple peer-reviewed randomized controlled trials and large-scale studies demonstrating reductions in pain and improvements in functional ability for musculoskeletal conditions. Similarly, Omada Health has published randomized controlled trials, such as the PREDICTS trial, showing significant and clinically meaningful improvements in HbA1c reduction and weight loss for diabetes prevention. Noom has also conducted randomized controlled trials, including a large-scale study published in June 2026, demonstrating sustained weight loss.
- Mental Health AI: Platforms like BetterHelp and Calm often rely on user-reported outcomes, satisfaction surveys, and internal efficacy studies. While BetterHelp participates in research and tracks outcomes, large-scale, independent randomized controlled trials specific to the platform are not widely reported. Similarly, for the Calm meditation app, evidence primarily stems from internal studies or general research on meditation, rather than platform-specific RCTs.
- Diagnostic AI (Beyond Cardiology): Even within regulated SaMDs, the depth of evidence varies significantly. Companies like iRhythm Technologies, with their Zio XT patch for arrhythmia detection, have accumulated substantial real-world evidence and observational data, and have also conducted randomized controlled trials, such as the GUARD-AF trial. However, the sheer volume and RCT rigor seen with HeartFlow are still rare across the broader diagnostic AI landscape.
This disparity highlights a critical gap. Payers, particularly those engaging in capitated or bundled payment models, are increasingly risk-averse. They need assurance that an AI solution will not only improve care but also demonstrably reduce downstream costs. Without the highest levels of evidence, these companies face an uphill battle in securing broad VBC adoption and favorable reimbursement terms.
CMS and CMMI: Shifting Expectations for Innovation
The Centers for Medicare & Medicaid Services (CMS) and its innovation arm, the Center for Medicare & Medicaid Innovation (CMMI), are pivotal in shaping the future of VBC. Their continued emphasis on outcomes-based payment models, such as ACO REACH and various bundled payment initiatives, directly elevates the demand for robust evidence from all participating technologies, including AI. The ACO REACH model is currently projected to conclude at the end of 2026, with CMS having announced a new model, the Long-Term Enhanced ACO Design (LEAD) Model, to run for ten years following its conclusion. CMMI’s ongoing evaluations of novel payment models consistently underscore the need for solutions that demonstrate clear clinical utility and financial performance. CMMI program evaluation methodology. The FDA’s SaMD framework and evolving guidance on Good Machine Learning Practice (GMLP) further solidify the regulatory landscape, pushing AI developers towards more rigorous validation. The GMLP principles were finalized by the International Medical Device Regulators Forum (IMDRF) in January 2025. While the FDA has issued various guidance documents, it also withdrew its guidance adopting IMDRF SaMD Clinical Evaluation principles in January 2026, signaling a move toward FDA-specific frameworks rather than direct incorporation of international standards. However, regulatory clearance for safety and efficacy (e.g., a 510(k)) does not automatically equate to VBC readiness. Payers require evidence of value, which often extends beyond the scope of initial regulatory approvals. The ISO 14155 standard for clinical investigation of medical devices provides a robust framework, but adherence to such standards for AI health platforms is not yet universal.
Conclusion: The Path to VBC Partnership is Paved with Evidence
For AI health platforms aspiring to meaningful participation in value-based care arrangements, the message is clear: the era of “trust us, it works” is over. Health plan executives and clinicians demand verifiable, peer-reviewed outcomes data, especially those demonstrating both clinical efficacy and financial performance. While many innovative AI solutions are making strides, only a select few have ascended the VBC Evidence Pyramid to its highest tiers. HeartFlow serves as a compelling case study, illustrating that rigorous scientific validation, culminating in extensive peer-reviewed RCTs, is not merely an academic exercise but a strategic imperative for securing VBC contracts and ultimately, transforming healthcare delivery. For those AI health companies currently operating at the lower tiers, investing in robust, externally validated research is no longer optional; it is the essential pathway to unlocking the full potential of AI in a value-driven healthcare ecosystem.
Frequently Asked Questions
What is the VBC Evidence Pyramid and why is it important for evaluating AI health platforms?
The VBC Evidence Pyramid is a five-tier framework for evaluating AI health platforms, ranging from pilot data to peer-reviewed Randomized Controlled Trials (RCTs). It is crucial because it provides a structured approach to discerning genuine impact from aspirational claims, ensuring AI solutions demonstrate measurable improvements in patient outcomes and financial performance with evidentiary rigor.
What level of evidence do health plans typically require for AI solutions in value-based care contracts?
For VBC contracts, health plans increasingly demand high-quality evidence, often seeking data from rigorous Real-World Evidence (RWE) studies with robust controls or, ideally, Peer-Reviewed Randomized Controlled Trials (RCTs). They require demonstrable, quantifiable improvements in validated cost reduction, clinical outcomes, and population health impact, rather than just promises of efficiency.
What specific outcomes data are payers looking for when considering AI health platforms for VBC contracts?
Payers are specifically looking for validated cost reduction (e.g., reduced ED visits, hospitalizations), clinical outcome improvement (e.g., better disease management, adherence to guidelines), and evidence of population health impact across diverse patient groups. They also require assurances of interoperability, data security, and regulatory clearance for AI tools classified as Software as a Medical Device (SaMD).
How does the VBC Evidence Pyramid help clinicians trust AI solutions?
The VBC Evidence Pyramid helps clinicians trust AI solutions by providing a clear hierarchy of evidentiary rigor, similar to established clinical evidence frameworks. By understanding where an AI solution’s evidence falls on this pyramid, clinicians can assess the reliability and generalizability of its claimed benefits, ensuring interventions are underpinned by robust evidence rather than just technological novelty.
