Work overview

Section 02 of 07

Methods

Evaluation of the SMS-ICU score for multicontinental use: A federated external validation study using critical care registry data

Alexander Tracy, Fathima Fazla, Jorge I.F Salluh, Rashan Haniffa, Rabiul Alam Md Erfan Uddin, Diptesh Aryal, Sylvia Brinkman, Gastón Burghi, Zakary Doherty, Dave Dongelmans, Stefano Finazzi, Ville Ihalainen, Bharath Kumar Tirupakuzhi Vijayaraghavan, Mohammed Basri Mat-Nor, Tahlia Perumal, David Pilcher, Luigi Pisani, Matti Reinikainen, Moses Siaw-Frimpong, Menbeu Sultan, David Thomson, Giovanni Tricella, Abigail Beane, and Otavio Ranzani · 2026

Contents

Section 02 of 07

  1. 01Introduction
  2. 02Methods
  3. 03Results
  4. 04Discussion
  5. 05Contributor roles
  6. 06Funding
  7. 07Declaration of competing interest
Text size
Work overview

Section 2 of 7

Methods

Alexander Tracy, Fathima Fazla, Jorge I.F Salluh, Rashan Haniffa, Rabiul Alam Md Erfan Uddin, Diptesh Aryal, Sylvia Brinkman, Gastón Burghi, Zakary Doherty, Dave Dongelmans, Stefano Finazzi, Ville Ihalainen, Bharath Kumar Tirupakuzhi Vijayaraghavan, Mohammed Basri Mat-Nor, Tahlia Perumal, David Pilcher, Luigi Pisani, Matti Reinikainen, Moses Siaw-Frimpong, Menbeu Sultan, David Thomson, Giovanni Tricella, Abigail Beane, and Otavio Ranzani · about 4 minutes

Study design, setting and participants

This retrospective federated external validation study was conducted using contemporaneous data from critical care registries on multiple continents.15 Registries were identified from membership of international collaborations, including Linking of Global Intensive Care (LOGIC) and Collaboration for Research Implementation and Training for Critical Care in Asia and Africa (CCAA).3,16,17 Additional registries known to LOGIC or the authors were invited to participate.

This study included patients admitted to ICUs during 2023. In Finland, patients admitted during 2022 were included instead due to internal data access restrictions. ICUs were eligible for inclusion if they started contributing data to their registry at least one year prior to the start of the study period. This provided a lag period of at least one year to maximise data quality.

As per the studies that developed and customised SMS-ICU, patients were excluded if they were: not over 18 years of age at ICU admission, without a known survival status for ICU or hospital admission, or transferred to another hospital.11,14 A sensitivity analysis excluded all interhospital transfers, both inward and outward. This accounted for the possibility that inward inter-ICU transfers may occur beyond the initial 24-h period during which SMS-ICU should be calculated.11 Unlike previous studies, elective ICU admissions were not excluded.

Central ethical approval was provided by the Alfred Hospital Ethics Committee (project no. 514/24) on 16th August 2024. Necessary local administrative approvals were obtained by each participating registry.

This study is reported according to TRIPOD guidelines.18

Federated approach

Analysis materials were prepared and distributed to participating registries. These included instructions for variable mapping, data dictionary, analysis scripts and study protocol (all available from https://github.com/NICST-PROTECT/sms-icu-validation/tree/main). The analysis scripts automatically applied study inclusion and exclusion criteria. Variable mapping, data processing and analysis were conducted internally by each registry, and aggregate results were reported without transfer of patient-level data.

Variables

The SMS-ICU component variables are age, blood pressure, admission type, haematological malignancy or metastatic cancer, vasopressor or inotrope use, respiratory support, and renal replacement therapy (Supplementary Table 1). SMS-ICU variables were defined according to the original study, as outlined in the data dictionary supplied to participating registries.11 The supplementary methods section describes any local adaptations to the original variable handling (Supplementary Methods).

To describe case-mix, registries were requested to report the reason for ICU admission using APACHE II diagnostic categories. However, alternative categorisations were permitted if harmonisation with APACHE II categories was not feasible. This maximised the range of registries able to participate.

The main outcome of interest was all-cause in-hospital mortality. Therefore, the primary analysis compared SMS-ICU predicted mortality to actual in-hospital mortality. Only registries that systematically collected hospital mortality outcomes were included in this analysis. In registries where only a subset of ICUs systematically reported hospital outcomes, the registries were filtered to include only ICUs with <20% hospital outcome missingness.

As all registries systematically collected ICU mortality outcomes, an a priori secondary analysis was performed comparing predicted mortality to actual ICU mortality. The purpose of this analysis was to test score performance in a broader range of settings, including those in which hospital outcomes are not yet systematically collected.

Sample size

Previous work has suggested that 100–200 outcome events are required to externally validate a predictive model.19,20 All eligible patients from each registry were included, and it was expected that this threshold would be met in all participating registries.

Statistical analysis

SMS-ICU scores (see Supplementary Table 1) were calculated.11 The customised equation for prediction of hospital mortality14 was used to predict outcomes for both primary and secondary analyses. All analyses were conducted in R version 4.4.

To assess discrimination, receiver operating characteristic (ROC) curves were plotted using the pROC package and the area under the curve (AUROC) was reported.21 The a priori threshold for acceptable discrimination was set at 0.7, in line with a commonly used definition.22 This threshold is expected to reflect adequate discrimination for international use; excellent discrimination is not required for this purpose. AUROCs from each registry were pooled using restricted maximum likelihood meta-analysis with the metafor package.23,24

Calibration curves were plotted with the rms package using both logistic and non-parametric LOWESS-smoothed calibration curves. Calibration-in-the-large was assessed using the calibration intercept. Both the calibration intercept and calibration slope were derived from the logistic calibration function. Observed/expected (O/E) ratios (equivalent to standardised mortality ratios) were reported for each registry.25

Overall model fit was assessed by calculation of Brier scores in each registry.26

Sensitivity analyses consisted of: (i) exclusion of both inward and outward interhospital transfers as described above, and (ii) a complete case analysis (see “Missing Data” below).

Subgroup analyses were performed for medical (non-surgical), emergency surgical, and elective surgical admissions. These subgroups were selected because SMS-ICU's applicability to elective cases was unknown. Adequate performance in each of these subgroups is a prerequisite for broad application across critical care populations.

A further subgroup analysis was conducted for patients receiving invasive mechanical ventilation. This aimed to assess score performance in a group with higher illness severity and risk of death.27

Forest plots comparing AUROC between subgroups were drawn using ggplot2.

Missing data

Missing data for SMS-ICU component variables were handled by modal imputation using R. Modal imputation was preferred to normal imputation because some non-scoring SMS-ICU categories, e.g. admission type, were expected to apply to a minority of patients. Multiple imputation was not used due to the expected low missingness rate, the low number of variables available for inclusion in a multiple imputation model, and the methodological complexity of implementing this in a federated study. A sensitivity analysis considered only complete cases with no imputation.

Missing outcomes were handled by excluding individuals for whom the outcome was missing.