Counting Who We Serve: How Federal Agencies and Tribal Leaders Built a New Way to Track Indian Health Coverage
The CMS–IHS Data Match Project, 2017–2023
A Lecture Paper for Introduction to Indian Health
Amended — claims-data definition clarified
Overview and Purpose
This paper introduces the CMS–IHS Data Match Project: what it is, why it matters, how it was built, and what it tells us about American Indian and Alaska Native health policy.
Section I: Why Does This Matter?
For years, Indian health financing research lacked a complete, verified count of how many AI/AN Medicaid enrollees were also registered with IHS. Without that administrative connection, it was difficult to measure the relationship between Medicaid and Indian health programs. The Data Match Project was designed to address that problem.
Key Actors
- Indian Health Service (IHS): Maintains registrant and health-program records.
- Centers for Medicare & Medicaid Services (CMS): Administers Medicare and Medicaid and maintains enrollment records.
- National Indian Health Board (NIHB): Served as the lead research organization.
- CMS–IHS Tribal Advisory Committee: Provided Tribal oversight and attention to data sovereignty.
Section II: The Old Way — Survey Data and Its Limits
Before the Data Match Project, researchers relied heavily on American Community Survey data. Survey data can estimate coverage patterns, but it cannot substitute for direct administrative records when the question is whether a person is actually enrolled or registered.
The distinction is fundamental: survey data provides an estimate; administrative data can provide a record of participation in a program.
Section III: Building the Match
Creating the match required planning, legal agreements, secure technical infrastructure, and sustained Tribal engagement. IHS registrant records and CMS Medicaid enrollment records were linked using identifiers common to the systems.
Important distinction about claims: In this paper, claims data is used only to describe a dichotomous indicator — whether a Medicaid-paid claim was present or absent for the matched person during the period. The analysis did not use the contents of individual claims. No information about the service, provider, diagnosis, procedure, or amount paid was used. Accordingly, the presence of a claim should not be interpreted as a measure of utilization, type of service, or expenditure.
Deterministic matching looks for exact agreements across identifiers. Probabilistic matching uses scoring methods to identify likely matches when small inconsistencies occur.
Data Sovereignty and Governance
Governance was central to the project. Secure data environments, data-use agreements, privacy protections, small-cell suppression, state-level reporting, and Tribal Advisory Committee oversight were part of the research design.
Section IV: What They Found
The Data Match produced a multi-year administrative dataset covering 2019–2023. The project identified more than 940,000 individuals in the matched IHS-registrant/Medicaid-enrollee population.
This is an administrative count rather than a survey estimate. It provides a more direct measure of the relationship between IHS registration and Medicaid enrollment.
From Enrollment to Dollars
Earlier analyses associated the matched enrollment population with estimated Medicaid expenditures using per-member-per-year (PMPY) benchmarks. These expenditure estimates should not be confused with claim-level expenditures from the Data Match itself.
What the Match Did Not Include
The claims measure was a dichotomous indicator: a Medicaid-paid claim was either present or absent. The analysis did not use information about the service, provider, diagnosis, procedure, or amount paid.
Therefore, the project should not be described as providing claim-level utilization or expenditure data. A paid-claim indicator establishes only that a Medicaid-paid claim occurred; it does not describe what happened on the claim or how much was paid.
Section V: Why Administrative Data Matters
Administrative data are generated as agencies operate their programs. They can provide direct information about enrollment and registration within those systems, although they can also contain errors and require careful interpretation.
NIHB’s leadership also illustrates the importance of Tribal-led research: federal agencies may hold the data, but Tribal participation can shape the research questions, governance, interpretation, and use of findings.
Section VI: Discussion Questions
- How does accurate health data connect to the federal trust responsibility?
- Why might a Tribal organization serve as the lead research organization?
- What are the tradeoffs between privacy protections and Tribal-level reporting?
- How might expenditure estimates based on PMPY benchmarks be used or misused when the underlying match does not contain claim-level payment information?
- What does the project’s multi-year development tell us about health-data research in Tribal contexts?
Section VII: Key Takeaways
- Health disparities require precise data to address.
- Tribal sovereignty extends to data.
- Federal-Tribal partnerships are complex but essential.
- Administrative data is powerful but not perfect. The match could distinguish beneficiaries with and without a paid claim, but it did not provide claim-level utilization or payment information for analysis.
- Tribal-led research can shape both the questions asked and the governance of the resulting evidence.
Glossary
- Administrative Data: Data generated by government agencies while carrying out their programs.
- Data Matching / Administrative Linkage: Combining records from separate databases to identify individuals appearing in both.
- Deterministic Matching: Matching based on exact agreement across specified identifiers.
- PMPY: Per-member-per-year, an average annual cost measure used in fiscal modeling.
- Probabilistic Matching: A method using scoring algorithms to identify likely matches despite minor inconsistencies.
- Tribal Data Sovereignty: Tribal authority over how data about Tribal communities are collected, used, stored, and shared.
- VRDC: A secure CMS data environment for approved researchers.