Section 1 of 9
Introduction
Qian Wang, Yu‐xiang Song, Xiao‐dong Yang, Jing‐wei Zhang, Rui‐zhe Sun, Xiao‐dong Wu, Ao Li, Jing‐sheng Lou, Hao Li, Yan‐hong Liu, Jun‐mei Xu, Di‐fen Wang, Qing‐ping Wu, Yu‐ming Peng, Yi‐qiang Chen, Jiang‐bei Cao, and Wei‐dong Mi · about 3 minutes
Older adults undergoing surgery are at increased risk of postoperative complications because of diminished physiological reserve, multimorbidity, frailty, and heightened vulnerability to perioperative stress. Among these complications, postoperative delirium (POD) and acute kidney injury (AKI) are two of the most frequent and clinically important adverse events [1, 2]. POD is associated with prolonged hospitalization, functional decline, loss of independence, cognitive deterioration, and increased mortality [3, 4, 5], whereas AKI is linked to longer hospital stay, greater resource utilization, progression to chronic kidney disease, and increased short‐ and long‐term mortality [6, 7]. Although POD and AKI differ in pathophysiology, clinical presentation, and perioperative management, both are highly relevant in older surgical patients and may be mitigated through timely identification of high‐risk individuals. Thus, accurate preoperative or early perioperative risk prediction for each complication is essential to support individualized prevention, monitoring, and resource allocation.
Over the past decade, numerous prediction models for POD or AKI have been developed using conventional statistical methods and, more recently, machine learning (ML) techniques [8, 9, 10]. These models typically integrate demographic characteristics, comorbidities, laboratory indices, medication exposure, and intraoperative variables to estimate perioperative risk [11, 12]. However, their clinical applicability remains limited. Many models were derived from single‐center cohorts or relatively homogeneous populations and lacked rigorous external validation [13, 14, 15]. Consequently, apparent strong performance in development datasets may not translate well across institutions with different patient case mixes, perioperative workflows, surgical profiles, and outcome ascertainment practices [16].
Developing robust multicenter prediction models for perioperative complications is further complicated by growing barriers to centralized data sharing. Regulatory requirements, institutional governance policies, privacy concerns, and heterogeneity in electronic health record systems often restrict direct pooling of patient‐level data across hospitals [17, 18, 19]. In perioperative research, these challenges are amplified by the need to integrate information spanning preoperative assessment, intraoperative management, laboratory testing, and postoperative surveillance, all of which may differ substantially across centers [20]. In addition, multicenter clinical data are often non‐independent and non‐identically distributed (non‐IID), reflecting differences in patient selection, surgical complexity, anesthetic practice, monitoring intensity, and documentation patterns [21, 22, 23]. Under such conditions, traditional centralized ML pipelines may be difficult to implement and may remain vulnerable to site‐specific bias, thereby limiting model transportability and real‐world utility [24].
Federated learning (FL) offers a decentralized paradigm for collaborative model development that allows participating centers to train a shared model without transferring raw patient data to a central repository [25]. By preserving data within institutional boundaries, FL can facilitate multicenter collaboration while reducing privacy, governance, and re‐identification risks [26, 27]. Prior studies have demonstrated the feasibility of FL in several medical domains, including neuroimaging, oncology, dermatology, and other risk stratification tasks [28, 29, 30, 31]. Compared with isolated local learning models (LLMs), FL may improve generalizability by leveraging broader institutional diversity, while compared with centralized learning models (CLMs), it offers a privacy‐preserving alternative that may be more feasible in real‐world healthcare settings. However, important challenges remain, including inter‐site heterogeneity, imbalanced sample sizes, and variation in local feature distributions and outcome prevalence. In the perioperative setting, multicenter FL applications remain scarce, and evidence is particularly limited for separate prediction of POD and AKI in older surgical populations.
To address these gaps, we conducted a multicenter study across five hospitals and retrospectively evaluated a simulated federated learning framework for patients aged ≥ 65 years undergoing non‐cardiac, non‐neurosurgical procedures. We developed and validated separate federated learning models (FLMs) for POD and AKI, and implemented three federated strategies: federated averaging (FedAvg), federated proximal optimization (FedProx), and federated local self‐distillation (FedLSD). We benchmarked these federated models against site‐specific LLMs and CLMs, and evaluated their discrimination, calibration, and clinical utility using decision curve analysis (DCA). In addition, we applied Shapley Additive Explanations (SHAP) to identify influential perioperative predictors and to explore both shared and center‐specific feature importance patterns across institutions and across the two outcomes. By separately modeling POD and AKI within a privacy‐preserving multicenter framework, this study aims to provide a scalable and interpretable approach for perioperative risk stratification in older adults and to support more targeted prevention and management strategies for these two major complications.