Key Takeaways
- Quality measurement fragmented across payers, accrediting bodies, and government agencies produced more than 400 endorsed NQF measures without solving the underlying problem, because fragmented governance produces fragmented output.
- Claims data were built for billing, not clinical measurement. When richer clinical data is applied, 27% to 94% of hospitals rated high-quality on administrative data get reclassified as intermediate or low quality.
- The translation gap is measurable. For smoking cessation, chart abstraction showed a 94% performance rate while EHR-generated reports on the same patients showed 11%. Same care, opposite conclusion.
- Care is fragmented too. 79% of patients use more than one facility in a year, so single-site measurement works with an incomplete picture.
- Concentration beats proliferation. AHRQ analysis found seven indicators drive 93% of measurable population health benefit. FHIR, USCDI, and TEFCA now let professional organizations author executable specifications payers implement directly.
Healthcare has spent decades and billions of dollars trying to measure its way to quality, and the results have underwhelmed. The ambition was sound. If we can quantify quality, we can reward it, and if we can reward it, we can accelerate it. Somewhere between the clinical guideline and the payer contract, something fundamental broke down. The organizations built to fix it, including NCQA, NQF, and AHRQ, have not been able to produce the transformation the system needs. The reasons are structural, not personal, and the emerging interoperability stack finally makes a different approach possible.
This analysis traces why the current system underdelivered, why the failure lives at the technical level of how measures get specified and handed off, and why professional medical organizations are uniquely positioned to close the gap. The argument runs from diagnosis to infrastructure to a strategic roadmap that specialty societies can act on now.
Why decades of quality measurement underdelivered
The original sin of American quality measurement is fragmentation. Professional organizations, payers, accrediting bodies, and government agencies each built their own quality ecosystems, often measuring the same conditions with subtly different specifications. The result is a landscape where the same provider can look high-quality on one scorecard and mediocre on another, not because care changed, but because definitions did.
The National Quality Forum was established specifically to prevent this proliferation. It now endorses more than 400 quality measures, and yet the quality problem remains elusive and in some respects compounded. This is not a criticism of the people involved. It is a structural critique. Fragmented governance produces fragmented output, and fragmented output produces administrative burden with unclear proportional clinical benefit.
Where NCQA was formed largely by purchasers seeking accountability, professional societies are formed by and for clinicians. Where NQF has historically accepted measures from multiple competing entities operating under different methodological standards, professional organizations hold the guideline ownership, specialty registry infrastructure, and clinician trust that integrated frameworks require. The failures of the current system are not evidence that quality measurement is impossible. They are evidence that measurement built without clinical ownership, without harmonization to practice guidelines, without adequate data, and without meaningful incentives produces bureaucracy rather than improvement.
The claims data dead end
Alongside fragmentation sits a second and more insidious problem, the persistent reliance on administrative claims data for clinical quality measurement. Claims were designed for billing. They capture what was done and what was charged. They do not capture why, how well, or with what clinical context.
When measure developers use claims as their primary substrate, they are forced to approximate clinical populations using diagnosis codes that lack the granularity of the original guideline. A guideline might define a population using ejection fraction, biomarker thresholds, or symptom severity scores. A claims-based measure must approximate that same population using ICD codes. The result is systematic misclassification driven by data asymmetry, which may sit at the root of payer-provider abrasion in many cases.
Studies have found that between 27% and 94% of hospitals classified as high-quality using administrative data were reclassified as intermediate or low quality once richer clinical data was applied.
For a system making payment and public reporting decisions on those classifications, a reclassification range that wide is not a rounding error. It is over-indexing on measures of perceived value, with consequence at scale. The streetlight effect applies cleanly here. We have been measuring where the light is, not necessarily where quality actually lives.
Burden without benefit
The downstream effect of proliferating measures built on flawed data is a system that imposes enormous cost on providers without clear return. Primary care practices carry the weight of thousands of electronic clinical quality measures developed by independent groups, many inconsistent with actual clinical guidelines. Clinicians and embedded quality auditors spend time on measure compliance that could be spent on care.
Analysis of AHRQ's quality indicators found that just seven of them account for 93% of total measurable population health benefit. The remaining hundreds of measures, each demanding documentation, abstraction, and reporting, account for the other 7%. This is not a portfolio problem. It is a misallocation crisis, and it points directly at the discipline the system has lacked.
Marginal incentives, marginal results
Most damning is that the programs built on this measurement infrastructure have not worked. Research on Medicare's Hospital Value-Based Purchasing program found no evidence of improved clinical process or patient experience measures in its first nine months, and no reduction in mortality in its first thirty months. Systematic reviews describe the impact of value-based purchasing broadly as mixed and far from convincing.
The critique is not that value-based payment is wrong in principle. It is that value-based payment layered on top of flawed measurement, inadequate data, and marginal incentives produces marginal results. Fix the substrate and the incentive has something real to attach to.
The translation gap
The misalignment is easiest to see at the technical level, in the handoff between the clinician who defines a measure and the actuary who implements it. When a clinical guideline committee defines a quality measure, it thinks in clinical language, identifying patient populations by lab values, imaging findings, symptom severity scales, and biomarker thresholds. This is also where the life sciences industry anchors its research and development. A heart failure measure might specify a population using ejection fraction below 40%. An oncology measure might require confirmed pathologic staging. These are the criteria that distinguish the patients for whom an intervention is appropriate.
When a payer actuary implements that measure, the reality is different. They have claims. Claims carry diagnosis codes, procedure codes, and encounter dates. They do not carry ejection fractions, staging results, or biomarker levels. So the actuary approximates using ICD codes that roughly correspond to the clinical population, making assumptions where clinical context is absent, and building logic that cannot fully represent what the guideline intended. This is the translation gap, and its consequences are profound.
One study of cardiovascular quality measures found that for smoking cessation, medical record abstraction suggested a 94% probability of meeting the performance target while EHR-generated reports on the same patient population suggested only 11%. Same patients. Same care. Opposite conclusions.
That is not a data problem. It is a structural problem with how quality measures are specified and handed off. When population identification is inaccurate, results are unreliable, and practices get penalized or rewarded based on measurement artifacts rather than actual care quality.
Why single-site data is not enough
The translation gap is compounded by the fragmentation of care delivery itself. Research shows that 79% of patients receive care at more than one facility during a calendar year. A quality measure calculated using only the data available to a single practice or health system is working with an incomplete picture by construction.
When health information exchange data was incorporated into quality calculations in published studies, 15% of all measure calculations changed, affecting nearly one in five patients. These were not edge cases. They were systematic errors introduced by incomplete data. Under the current model, neither the clinician nor the payer has visibility into the full picture, and both make consequential decisions anyway.
The infrastructure that changes the equation
The healthcare system now has building blocks that make a different approach not just theoretically possible but practically achievable. Three standards stack together to carry a quality measure from clinical intent to accurate execution.
CMS has designated FHIR as the standard for interoperability and now requires payers to establish patient access APIs using it. The Interoperability and Prior Authorization Final Rule creates bidirectional data flow, enabling payers to access clinical data for quality measurement while returning claims data to clinicians. The regulatory momentum has arrived, and it points in exactly the direction the measurement problem requires.
Proof the model works
This is not a theoretical future. ASCO, ASTRO, and their partners have already demonstrated feasibility through the CodeX Quality Measures for Cancer project, using FHIR standards and specialty-specific data profiles known as mCODE to author, test, and execute oncology quality measures. Their findings were direct. Measures generated accurate results when executed against FHIR repositories, burden was reduced for all parties in the quality measure lifecycle, and the approach created a less burdensome path for everyone involved.
The technology works. The infrastructure is coming. The regulatory mandate is in place. What is missing is specialty-level clinical ownership of quality measure specifications built for this new environment.
From narrative measures to executable specifications
The shift this enables is not incremental. Under the paradigm that has governed quality for decades, professional organizations develop quality metrics in narrative form, payers interpret those narratives using claims data, interpretation gaps are inevitable, measurement accuracy suffers, and clinicians bear the burden of a system that does not reflect their actual practice. The emerging paradigm inverts the control point. Professional organizations develop FHIR-based quality measure specifications as executable code that payers implement directly using clinical data accessed through TEFCA. The measure travels with its clinical logic intact, population identification uses the criteria the guideline intended, calculation is automated and auditable, and interpretation gaps are eliminated by design.
That is a structural shift in who controls the quality specification and how accurately care actually gets measured. It also asks professional organizations to expand their conception of what measure development means. Writing a narrative specification is no longer enough. The new standard is developing FHIR implementation guides, authoring Clinical Quality Language specifications, mapping specialty-specific data elements to USCDI, and partnering with payers on implementation through TEFCA-enabled exchange. This requires technical infrastructure, informatics capability, and a genuine commitment to clinical data stewardship well beyond guideline publication.
The organizations that build this capability first will define quality measurement for their specialties for the next generation and defend the posture of the professional community they serve. Those that do not will find their guideline-based measures continually approximated, and at times misrepresented, by payers working with the only tools available to them, and the abrasion will persist in some form.
Integrate economic value into guidelines
The most powerful thing a professional organization can do to align with payer value frameworks is to speak the language payers actually use, which is cost-effectiveness. This does not mean subordinating clinical judgment to economic metrics. It means acknowledging that health systems and payers are making allocation decisions whether or not clinical guidelines inform them, and that specialty societies with the deepest clinical evidence base are uniquely positioned to provide that guidance credibly.
The model already exists in practice. The ACC/AHA 2025 Guideline Methodology Manual explicitly incorporates cost and value analysis as a formal component of clinical practice guideline development. Cost-effectiveness is embedded in the methodology from the outset, with standardized thresholds, structured value classifications, and a framework for rating the certainty of economic evidence alongside clinical evidence. That is a meaningful structural commitment, and it sets a bar other specialty societies should study and build toward.
The American College of Physicians has taken a complementary position on the measurement side, arguing publicly for a smaller number of high-impact, evidence-grounded conditions on the grounds that not everything that can be measured should be. Their call for a unified national focus on the most important and common clinical conditions is exactly the kind of institutional position that can move the conversation at the federal and payer level. What this requires is a methodological commitment to treat economic evidence with the same rigor as clinical evidence, to establish reference standards for future analyses, and to produce transparent value classifications even when the results are inconvenient. It aligns cleanly with CMS approaches around the health economics and outcomes research evidence required of industry to attain new billing codes.
Couple economic value with recommendation strength
One of the most transformative moves available to professional organizations is to link economic value directly to recommendation classification, so the strength of a clinical recommendation reflects not just clinical evidence but cost-effectiveness evidence as well. This also creates research anchor points where known unknowns are cleanly evaluated and set up for active research and validation.
It creates something the current system almost entirely lacks, which is market incentives embedded in guideline structure. When a Class I recommendation requires both strong clinical evidence and cost-effectiveness at an accepted threshold, it changes what manufacturers, health systems, and payers do. Manufacturers gain an incentive to price interventions in alignment with the value they generate. Payers gain a credible, clinically grounded basis for coverage decisions. Clinicians gain guidance that integrates both the science of benefit and the reality of resource allocation. This is transformational rather than incremental. It moves professional organizations from publishing guidelines that payers interpret inconsistently to shaping the economic architecture of care delivery in their specialty.
Focus on the vital few
The clearest lesson from two decades of measurement failure is that proliferation is not a strategy. The AHRQ finding that seven indicators drive the overwhelming majority of measurable population health benefit should function as a design principle, not a footnote. Measurement programs have kept expanding anyway, imposing burden without proportional value.
Professional organizations should resist the temptation to be comprehensive. The goal is not to measure everything but to measure what matters, measure it accurately, and create incentives that drive meaningful change. In practice that means actively retiring low-value process measures that do not link to improved outcomes, concentrating development resources on the vital few with demonstrated population health impact, preferring outcome measures over process proxies, and being willing to say publicly and credibly that measuring certain things is not worth the administrative cost of doing so.
Build toward equity from the start
Value-based payment has enormous potential as a lever for addressing health equity, and equal potential to entrench disparities if equity is not built into program design from the outset. The difference is architectural, and it has to be decided early.
Professional organizations should embed health equity considerations into both guideline development and quality measure design, incorporating differential effectiveness across populations into cost-effectiveness frameworks, structuring data elements to enable equity analysis, adjusting for social determinants of health in risk models, and including patient priorities beyond clinical endpoints in how quality gets defined. This is both the right thing and the strategically smart thing. Payers and CMS are increasingly focused on health equity as a dimension of quality, and specialty societies that develop credible, clinically grounded equity frameworks will find receptive audiences among exactly the policy and payer stakeholders they need to influence.
Engage payers as technical partners
The old model of professional organization and payer interaction is largely adversarial or transactional. Societies publish guidelines, payers develop coverage policies, gaps emerge, and utilization management and appeals processes fill in the cracks. The new model has to be different, because FHIR-based quality measurement only succeeds if payers can implement it.
That requires professional organizations to do something unfamiliar, which is to sit with payers in technical implementation conversations, pilot clinical data-based quality calculation in real contracting environments, and create feedback loops between implementation experience and specification refinement. This is not about giving payers control over clinical standards. It is about ensuring that the specifications professional organizations develop are actually implementable, and that when implementation barriers emerge, and they will, they get resolved through clinical input rather than administrative approximation.
A three-phase strategic roadmap
The transition does not happen in a single move. It sequences across roughly five years, with each phase building the capability the next one depends on.
Phase 1, Foundation (Years 1 to 2). Establish quality measurement infrastructure grounded in clinical data rather than claims. Critically assess existing measure portfolios and retire low-value indicators. Begin FHIR profile development by mapping specialty-specific clinical data elements to FHIR resources, leveraging and stewarding Value-Set Authority Center frameworks. Establish connections with TEFCA-qualified health information networks.
Phase 2, Integration (Years 2 to 4). Incorporate cost-effectiveness systematically into guideline development. Author FHIR-based quality measures using Clinical Quality Language. Pilot clinical data-based quality calculation with early-adopter payers. Develop payer-ready implementation guides. Generate economic evidence for priority clinical areas.
Phase 3, Transformation (Years 4 to 5 and beyond). Couple economic value with recommendation strength across the guideline portfolio. Scale FHIR-based measurement across payer contracts. Advocate for CMS adoption of FHIR-based measures in federal value-based programs. Establish continuous improvement feedback loops between clinical data, quality measurement, and guideline updates.
What is actually at stake
The quality measurement system did not fail because of bad intentions. It failed because of structural misalignment, with measurement authority held by bodies without clinical ownership, data systems inadequate to the specifications being enforced, and incentives too small to drive meaningful change.
Professional organizations hold the clinical authority, the registry infrastructure, the guideline ownership, and the stakeholder trust to build something fundamentally different. They also have every incentive to do so, because the alternative is continued delegation of clinical quality definition to organizations and data sources that cannot represent what the specialty actually knows about good care. The technology is ready, the regulatory environment is aligning, and the infrastructure is being built.
The paradigm shift is clear enough. The only question is which professional organizations will lead, and which will be led.
References
- National Quality Forum. Consensus-endorsed quality measures portfolio. qualityforum.org
- Agency for Healthcare Research and Quality. AHRQ Quality Indicators. Analysis cited for the concentration of measurable population health benefit across a small set of indicators. qualityindicators.ahrq.gov
- Peer-reviewed comparisons of hospital quality classification using administrative versus clinical data, reporting reclassification of 27% to 94% of hospitals rated high-quality on administrative data once clinical data was applied. Summarized in the source analysis.
- Published comparison of measurement methods for cardiovascular quality measures, reporting a 94% smoking-cessation performance rate by medical record abstraction versus 11% by EHR-generated report on the same population.
- Published studies of multi-facility care and health information exchange in quality measurement, reporting that 79% of patients use more than one facility per year and that incorporating HIE data changed 15% of measure calculations.
- Studies of Medicare's Hospital Value-Based Purchasing program and systematic reviews of value-based purchasing describing mixed and far from convincing effects on process, experience, and mortality outcomes.
- American College of Cardiology and American Heart Association. 2025 Guideline Methodology Manual, incorporating cost and value analysis into clinical practice guideline development. professional.heart.org
- American College of Physicians. Position urging a unified national focus on important and common clinical conditions in quality measurement. acponline.org
- HL7 CodeX. Quality Measures for Cancer project using FHIR and mCODE, with ASCO and ASTRO, demonstrating accurate execution against FHIR repositories and reduced burden. codex.hl7.org
- Centers for Medicare and Medicaid Services. Interoperability and Prior Authorization Final Rule (CMS-0057-F), establishing FHIR-based patient access APIs and bidirectional data flow. cms.gov
All views, analyses, and frameworks presented here reflect independent professional judgment informed by more than two decades of experience across payer strategy, clinical transformation, and health system operations. They do not represent the views or positions of any current or former employer or affiliated organization.