Work overview

Section 04 of 07

Discussion

Evaluation of the SMS-ICU score for multicontinental use: A federated external validation study using critical care registry data

Alexander Tracy, Fathima Fazla, Jorge I.F Salluh, Rashan Haniffa, Rabiul Alam Md Erfan Uddin, Diptesh Aryal, Sylvia Brinkman, Gastón Burghi, Zakary Doherty, Dave Dongelmans, Stefano Finazzi, Ville Ihalainen, Bharath Kumar Tirupakuzhi Vijayaraghavan, Mohammed Basri Mat-Nor, Tahlia Perumal, David Pilcher, Luigi Pisani, Matti Reinikainen, Moses Siaw-Frimpong, Menbeu Sultan, David Thomson, Giovanni Tricella, Abigail Beane, and Otavio Ranzani · 2026

Contents

Section 04 of 07

  1. 01Introduction
  2. 02Methods
  3. 03Results
  4. 04Discussion
  5. 05Contributor roles
  6. 06Funding
  7. 07Declaration of competing interest
Text size
Work overview

Section 4 of 7

Discussion

Alexander Tracy, Fathima Fazla, Jorge I.F Salluh, Rashan Haniffa, Rabiul Alam Md Erfan Uddin, Diptesh Aryal, Sylvia Brinkman, Gastón Burghi, Zakary Doherty, Dave Dongelmans, Stefano Finazzi, Ville Ihalainen, Bharath Kumar Tirupakuzhi Vijayaraghavan, Mohammed Basri Mat-Nor, Tahlia Perumal, David Pilcher, Luigi Pisani, Matti Reinikainen, Moses Siaw-Frimpong, Menbeu Sultan, David Thomson, Giovanni Tricella, Abigail Beane, and Otavio Ranzani · about 4 minutes

Main findings and implications

This study aimed to evaluate the suitability of SMS-ICU for use in critical care registries across multiple continents. Importantly, the prevalence of missing data for SMS-ICU variables was mostly low, suggesting that application to a diverse range of settings is feasible. This is because SMS-ICU uses a subset of variables already widely collected by registries.

Overall, discrimination was acceptable, with a pooled AUROC of 0.76 (0.72–0.80) in the primary analysis. This suggests that SMS-ICU could be appropriately used as an index of illness severity in international critical care populations. This could strengthen collaborative international research and facilitate the synthesis of international results in systematic reviews and meta-analyses. Better discrimination can be achieved using additional variables or more complex models;28,29 however, use of such models across a diverse range of settings would not be feasible.

The secondary analysis revealed that discrimination was poorer in African and Asian settings (Table 4). The observed discriminatory performance is consistent with previous studies of other scores.7,30 One explanation for this finding could be differences in case-mix between settings. Alternatively, in some contexts, the first 24 h of an ICU admission may be less informative regarding a patient's overall prognosis. For example, this may be the case if ICU admission typically occurs later in a patient's illness or follows marked variation in preceding healthcare. In some contexts, there may also be significant heterogeneity, e.g. in socioeconomic or nutritional factors, that SMS-ICU does not capture but that nonetheless influences prognosis.

Calibration, using the model derived by Zampieri et al.,14 varied substantially between registries. In general, the model overestimated mortality in high-income countries while underestimating mortality in lower-income settings. SMS-ICU was well-calibrated in Brazil and Uruguay, reflecting the location of its customisation.14 Calibration and model fit were particularly poor in Bangladesh, Ethiopia and Ghana. However, these registries had relatively small sample sizes, so these results should be interpreted with caution.

Possible reasons for the observed variation in calibration include inter-registry differences in case-mix, ICU performance, healthcare resourcing and health system characteristics. International variation in model calibration is well recognised, having previously motivated the development of regional customisations of the Simplified Acute Physiology Score (SAPS).31 Such variation contrasts with the excellent calibration that can be achieved by location-specific prognostic models.32,33

Score performance in subgroup analyses was typically slightly poorer than in the overall population. Because subgroups were defined by SMS-ICU variables, each subgroup analysis effectively removed one component from the score, thus reducing the information available to the model. In addition, the invasively ventilated subgroup is expected to have higher illness severity than the broader study population, making discrimination between individuals inherently more challenging. Therefore, additional variables are likely needed to discriminate between these patients.

Notably, prior to this study, SMS-ICU had not been validated for elective ICU admissions. This study shows that discriminatory performance in elective surgical cases is comparable to that for acute admissions. Therefore, SMS-ICU is applicable to a broader patient population than originally intended.

Taken together, these findings clarify the suitability of SMS-ICU for different international use cases. Given that discriminatory performance was acceptable, SMS-ICU can reasonably be used to describe illness severity. However, international variation in calibration limits its suitability for applications that require reliable absolute risk prediction. For example, these data do not support the routine use of SMS-ICU to estimate expected mortality rates for ICUs across heterogeneous settings.

Strengths and limitations

This study describes the performance of SMS-ICU in a large, diverse, multicontinental population of consecutive patients from critical care registries. This was enabled by its federated design, which obviated the need for international transfer of sensitive patient-level data. There is increasing interest in using such approaches to exploit large real-world datasets while maximising data security.[34], [35], [36] Building on this, the present study demonstrates the value of federated analysis for external validation of a prognostic model.

The handling of some variables required minor adjustments in certain registries to enable the calculation of SMS-ICU scores. This may have adversely affected performance in the affected registries. However, SMS-ICU performed relatively well in settings where this was required.

Although the use of registry datasets enabled real-world evaluation of score performance, there are limitations associated with this data source. Many registries lack full national coverage, so not all are nationally representative.15 Individuals admitted to ICUs for palliative care were not excluded, as not all registries identify such patients. Some registries had few contributing ICUs, limiting statistical power and generalisability of the results. This is a particularly important caveat to the findings in African registries. Finally, not all geographical regions were represented in this study.

Conclusion and future directions

In summary, SMS-ICU discrimination and model fit were broadly acceptable within many registries. This suggests that SMS-ICU could reasonably be used as an international index of illness severity. Variation in calibration limits other international uses of SMS-ICU in its current form.

In future, this variation could be addressed by customisation to re-calibrate the model based on contextual variables, such as geographical location, economic indicators or healthcare system characteristics. This could be achieved by developing context-specific equations or by deriving a new multi-level model. However, such work would require the systematic collection of hospital outcomes by a wider range of critical care registries.