Work overview

Section 02 of 05

METHODS

Validity and Predictors of Medical Claims Coding for the Identification of Patients with Obesity

Effie L. Kuti, Emma Richard, Kevin Schott, Christopher L. Crowe, Vincent Willey, and Bonnie Donato · 2026

Contents

Section 02 of 05

  1. 01INTRODUCTION
  2. 02METHODS
  3. 03RESULTS
  4. 04DISCUSSION
  5. 05CONCLUSION
Text size
Work overview

Section 2 of 5

METHODS

Effie L. Kuti, Emma Richard, Kevin Schott, Christopher L. Crowe, Vincent Willey, and Bonnie Donato · about 5 minutes

Study Design and Data Source

This retrospective cohort study used data from the Healthcare Integrated Research Database (HIRD®), a large, geographically diverse healthcare database containing administrative claims for over 80 million people from commercial and Medicare Advantage/Supplemental health insurance plans in the Northeastern, Southern, Midwestern, and Western regions of the US since 1 January 2006. Clinical data from integrated electronic health records (EHR), including BMI, were available for 12% of patients in the HIRD over the period of interest. This study identified patients with a BMI ≥30 kg/m2 (ie, with obesity per the World Health Organization and the Centers for Disease Control) during the intake period, defined as 1 June 2021 to 30 October 2022.2,9 The intake period was selected to maximize use of the most contemporary data available at the time of analysis and to capture a period in which novel pharmacologic treatments for obesity management were in use in clinical practice.8 The index date was defined as the first observed BMI during the intake period, and therefore the date varied by patient. The first BMI measurement was used to anchor the analysis at the earliest available documentation of obesity status in the EHR, helping to avoid using future information and ensuring that baseline characteristics, coding assessments, and follow-up periods were properly aligned over time. Two baseline periods were employed in this study: A 6-month pre-index period, defined as the 6 months prior to the index date, and a full pre-index period, defined as any time prior to the index date going back to the earliest data available data using ICD-10-CM diagnoses codes (1 October 2015). The 6-month pre-index period was used to ensure the capture of medication use and healthcare encounters most proximal to identification of obesity. The full pre-index period was used to ensure the comprehensive capture of chronic conditions.

Two time periods were also employed to explore the presence or absence of ICD-10-CM coding for obesity among patients identified with BMI ≥30 kg/m2 This broader window was selected to reflect real-world variability in coding practices and healthcare utilization, recognizing that diagnosis codes may not be recorded at the same time as BMI measurement but rather during nearby clinical encounters. In contrast, a narrower 60-day window before and after the index date was used for validity calculations (sensitivity, specificity, positive predictive value [PPV], and negative predictive value [NPV]) to more closely align coding with the timing of BMI documentation in the EHR. This shorter window was intended to reduce temporal misalignment between measured obesity status and coding. However, because both windows include time after the index date, there is potential for temporal ambiguity, as some diagnosis codes may reflect documentation occurring after the BMI measurement rather than contemporaneously. These design choices represent a balance between capturing real-world coding practices and maintaining temporal proximity for validation.

Study Population and Cohorts

The study population consisted of adult patients with at least 1 BMI measurement available in the HIRD during the intake period. Eligible patients were continuously enrolled in medical and pharmacy benefits for at least 6 months prior to the index date. Patients with an index BMI ≥30 kg/m2 were stratified into 2 subgroups, patients with and without an ICD-10-CM code for obesity in medical claims in the 6 months prior to or after the index date. Obesity was identified using ICD-10-CM diagnosis codes E66.0, E66.01, E66.09, E66.1, E66.2, E66.8, and E66.9, along with BMI Z-codes Z68.3-Z68.4. Female patients with diagnoses indicative of pregnancy within 9 months prior to or after the index date were excluded from analyses. Patients with BMI <30 kg/m2 were included solely to enable calculation of specificity and NPV but were excluded from regression analyses focused on predictors of coding among patients with obesity.

Variables and Analyses

Demographics were evaluated as of the index date. Baseline comorbidities were assessed over the full pre-index period by the presence of ≥1 medical claim with a diagnosis for the comorbidity of interest. Baseline medication use and healthcare encounters were assessed over the 6-month pre-index period by the presence of ≥1 pharmacy claim for the medication of interest and ≥1 medical claim for health service of interest, respectively. Missingness was minimal for most variables given the use of administrative claims data. Race and ethnicity were derived using a proprietary multi-source algorithm with high completeness. For variables with missing values, missingness was retained as a separate category, and no imputation was performed.

To validate ICD-10-CM codes for obesity, sensitivity, specificity, PPV, and NPV were calculated. Each eligible patient was categorized into one of four categories: true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). A patient was considered a TP if their index BMI was ≥30 kg/m2 and they had ≥1 ICD-10-CM diagnosis codes for obesity around the index date (ie, 60 days prior to or after index date). A FP was defined as a patient with an index BMI <30 kg/m2, who had ≥1 ICD-10-CM diagnosis codes for obesity within the designated window. TNs, on the other hand, were patients whose BMI was <30 kg/m2 and had no ICD-10-CM diagnosis codes for obesity within the designated window. Finally, FNs were defined as those patients whose BMI was ≥30 kg/m2 but who had no ICD-10-CM diagnosis codes for obesity within the designated window. Sensitivity analyses were not performed and represent an area for future research.

The prevalence of ICD-10-CM coding for obesity was examined among all eligible patients with an index BMI ≥30 kg/m2 (ie, with obesity), as well as within subgroups characterized by different clinical characteristics (eg, with and without hypertension over the 6-month pre-index period). The denominator for these calculations was the total number of eligible patients with a BMI ≥30 kg/m2 and characteristic of interest. In contrast, the numerator comprised only those patients who also had a code for obesity within the designated window around the index date (ie, 6 months prior to or after index date).

Multivariable logistic regression was implemented to explore predictors of the presence or absence of ICD-10-CM codes for obesity within this designated window among patients with an index BMI ≥30 kg/m2. Candidate variables were selected a priori based on clinical relevance and prior literature, as well as descriptive analyses conducted within the study population. These variables included demographic characteristics (eg, age, sex, race/ethnicity, insurance type, and socioeconomic status), clinical characteristics (eg, BMI category, comorbid conditions), medication use, and healthcare utilization measures. To reduce model overfitting and improve interpretability, a backward elimination approach was applied, with variables retained in the final model based on statistical significance (P<.05) and clinical relevance. Multicollinearity among covariates was assessed using variance inflation factors, with a threshold of >10 indicating potential collinearity concerns. Model performance was evaluated using measures of discrimination and calibration. Discrimination was assessed using the C statistic (area under the receiver operating characteristic curve), while calibration was evaluated using the Hosmer–Lemeshow goodness-of-fit test. These diagnostics were used to ensure adequate model fit and stability of estimates. All covariates were measured during the pre-index period.