Work overview

Section 02 of 05

METHODS

Identifying High-Priority Profiles for Edentulism and Other Health Outcomes Through Population-Based Clustering

Marjorie A. Rosenberg · 2026

Contents

Section 02 of 05

  1. 01INTRODUCTION
  2. 02METHODS
  3. 03RESULTS
  4. 04DISCUSSION
  5. 05CONCLUSIONS
Text size
Work overview

Section 2 of 5

METHODS

Marjorie A. Rosenberg · about 5 minutes

Study Sample

The Medical Expenditure Panel Survey (MEPS), an annual survey that is representative of the U.S. civilian, non-institutionalized population, was used. The data are publicly available and were exempt from IRB review.15 The data used for this study are representative of individuals aged ≥18 years.

Individuals in Panel 24, a longitudinal cohort, were followed for 4 years (2019–2022 inclusive) with 9 rounds of interviews. Rounds 1 and 2 were in 2019, Round 3 could be either in 2019 or in 2020, Round 4 was in 2020, Round 5 was in either 2020 or 2021, Round 6 was in 2021, Round 7 was in either 2021 or 2022, and Rounds 8 and 9 were in 2022.

Critical to the analysis is the indication of whether a person is completely edentulous, a variable that is not used in the formation of the clusters. A binary variable was defined, indicating whether a person is edentulous at any time across the 4 years of the study, using the question Have you lost all of your upper and lower natural (permanent) teeth? from Rounds 3, 5, 7, and 9 for those aged ≥18 years at the beginning of the study. If a person answered yes to any of these 4 questions, they were considered completely edentulous. If a person answered no to all 4 questions, they were not considered completely edentulous. Only individuals who have answered at least 1 of the edentulous questions were included in this study.

In this study, the baseline risk was established with 8 variables to develop the clusters that are common features seen in the literature relating to edentulism. These variables, used as binary variables, were chosen close to the beginning of the study. Age (18–49 and ≥50 years), gender (male, female), perceived health status (PHS) (excellent/very good; good/fair/poor), mental health status (MHS) (excellent/very good or good/fair/poor), an indicator of unemployment for the entire round, and an indicator of any limitation (work, school, or home) were obtained from the first round of data. Smoking status (smoker, nonsmoker) was obtained from the third round. The income category relative to the federal poverty line (high/medium, low/near, poor/poor) was used from the first calendar year. PHS, MHS, and an indicator of any limitation together can proxy one’s perceived overall health and substitute for diagnosed medical conditions. Using these variables mitigates the need to see a provider and represents how individuals feel overall. Income has been used as a proxy for education as well as for the affordability and access to health insurance and health care.16,17 The goal was to determine whether these variables that define the baseline risk are associated with complete edentulism, death, and other external variables.

Although the data are longitudinal, the study methods do not reflect time-varying variables. As mentioned, edentulism was defined by aggregating 4 different measurements into 1 binary variable that indicates whether or not a person is edentulous. Similarly, death was measured by an indicator variable as to whether or not a person died in the 4-year period. The baseline risk was compared with these 2 variables as well as with other external variables (diabetes, CVD) (combining diagnosed chronic heart disease, angina, myocardial infarction, stroke, and other heart issues, and high blood pressure), and high psychological distress using the Kessler-6 (K6) score in round 8). Diabetes and CVD are measured similarly to edentulousness and death, as being diagnosed in any of the four years of the MEPS study. In this way all of the external variables are a reflection of a person’s cumulative medical history. The Kessler-6 score, measured at the end of the study, captures the state of distress at that time.

Inherent in a longitudinal study is missing data; however, the complex survey design adjusts for data loss to follow-up or non-response. If and when individuals died, left their reporting unit, became institutionalized, or became active-duty military, their future responses were set to “inapplicable”. As the baseline risk is set at the beginning of the study, these events do not impact the analysis. The construction of the edentulism variable results in no missing data, as only those who answer yes or no are included in the study. Observations (unweighted) missing the baseline risk variables are 7 for activity limitations, 1 for perceived health status and 13 for smoking. Responses are imputed for all but one observation based on a future round of data. One observation was imputed as a non-smoker, so as not to remove their observation from the study. The entire cohort has 5,565 observations with 4,402 observations 18 and older. The final sample contains 4,245 observations representing 246,190,382 persons 18 and older (99% of the weighted 18+ US population).

Instead of analyzing the variables as separate inputs, clusters of similar baseline risk are created. Clusters are an effective way to uncover latent patterns and relationships, in contrast to regression methods. A weighted method of k-modes (wKModes in R) designed for inputs that are categorical was used. This method accounts for survey weighting and chooses the mode of each category on the basis of the sum of the weights. Frequency-based weighting was used, in contrast to simple matching, to examine the rarity of a category. If individuals shared a rare trait, they were considered more similar than if they shared a common trait. The algorithm was set with random starts and with convergence set at a maximum of 100 iterations, as is current practice. Finally, the algorithm was run 40 times, and the number of clusters was chosen on the basis of the lowest sum of squares.18

Seventeen clusters were chosen on the basis of the results from an elbow plot, a commonly used method.19,20 A larger number of clusters was chosen to allow for those within a cluster to be more similar, while creating clusters that differ more from one another.21 Finally, clusters were labeled by ranking the percentage of edentulism within the cluster from 1 (highest) to 17 (lowest).

To better describe and visualize the relationships among clusters, a second-stage clustering was performed to connect the clusters using hierarchical clustering with complete linkage.22 The results are represented by a dendrogram that visually resembles an inverted tree. The dendrogram is used to reduce the number of clusters into 3 larger blocks of clusters. In this way, the relationships among the clusters can be visualized, where the underlying baseline risk variables in the individual clusters are used to decide on the number of blocks.

Clusters that have a similar baseline risk were investigated and validated by comparing differences among the clusters by both the edentulous and death rates (and other external variables). Because the input variables were categorical, the distance measure chosen was cosine similarity.23 Cosine similarity measures pairwise correlation and has been used in text classification, genetics, and principal components.24, 25, 26, 27