Section 2 of 5
Materials and methods
Yuhang Wang, Dandan Dong, Shengming Shi, Yupeng Wu, Apekshya Singh, Jiayi Xie, Qiuyang Chen, Jianwei Zhu, and Xiaofu Li · about 6 minutes
Study design and ethics
This dual-center retrospective study was approved by the Institutional Review Boards of the Second Affiliated Hospital of Harbin Medical University (YJSKY2025-086) and Zhuhai People’s Hospital (2025-66); informed consent was waived. We enrolled 312 patients with pathologically confirmed rectal adenocarcinoma: 246 from Center 1 (training/internal validation ratio 7:3) and 66 from Center 2 (external testing). Eligibility criteria mandated: (1) surgical pathology-confirmed rectal malignancy; (2) preoperative MRI examination ≤ 14 days before resection; (3) Preoperative serum carcinoembryonic antigen (CEA) and carbohydrate antigen 19-9 (CA19-9) were assessed within 2 weeks. Exclusion parameters involved: (1) history of neoadjuvant treatment; (2) incomplete PNI histopathology records; (3) non-diagnostic MRI artifacts; (4) mucinous adenocarcinoma subtypes; (5) MRI-to-surgery interval > 14 days. Center 1 cases were stratified into training (n = 172) and internal validation (n = 74) cohorts at a 7:3 ratio. Conversely, the entire Center 2 cohort (n = 66) comprised the external testing cohort (Fig. 1).

Fig. 1: A schematic overview of the selection algorithm, detailing both inclusion thresholds and exclusionary conditions applied in the multicenter research cohort
Prior to surgical intervention, all enrolled participants were subjected to MRI scans, with technical specifications and imaging workflows comprehensively described in Appendix 1. Preoperative predictors encompassed demographic variables, serum biomarkers (CEA ≥ 5 ng/mL; CA19-9 ≥ 37 U/mL), and MRI-derived parameters (tumor longitudinal diameter, mrT/N staging, mrEMVI/MRF status), systematically retrieved from institutional EHR and PACS databases. The PNI status for all patients was derived from the final postoperative histopathological reports. Both participating institutions follow standardized diagnostic protocols for rectal cancer. PNI was defined according to established criteria [5, 16] as: (1) tumor cells invading any of the three layers of the nerve sheath (epineurium, perineurium, or endoneurium); or (2) tumor cells in direct contact with the nerve and wrapping around ≥ 33% of its circumference. To ensure diagnostic accuracy, S100 immunohistochemical staining was routinely utilized in clinical practice to confirm nerve structures when H&E staining was inconclusive. These clinical reports represent consensus diagnoses reviewed by senior pathologists, serving as the reliable ground truth for this study.
Image collection and pretreatment
To optimize clarity, we implemented a generative adversarial network (GAN)-enhanced 4× super-resolution reconstruction framework via the OnekeyAI platform (https://github.com/OnekeyAI-Platform/onekey). Initial preprocessing incorporated noise reduction and intensity normalization across the MRI dataset. Subsequently, low-resolution analogs were generated through downsampling, paired with native high-resolution images for adversarial training to capture high-frequency anatomical details. The GAN structure integrated a resolution-enhancing generator and a discriminator evaluating reconstruction fidelity, iteratively refining image quality through competitive optimization (Supplementary Fig. 1). To mitigate “hallucinations” and ensure feature stability, we enforced rigorous quality control, including visual inspection by radiologists and quantitative verification via external validation.
Tumor boundaries were independently annotated by two board-certified radiologists via ITK-SNAP on the super-resolution (SR) reconstructed images, with discordant interpretations adjudicated by a third expert (> 20 years’ experience) to ensure segmentation consensus. All images underwent spatial standardization (1 mm³ isotropic voxels) using a fixed-resolution protocol, ensuring cross-modal uniformity and model robustness. Radiomic feature stability was quantified through interobserver correlation coefficient (ICC) analysis on 30 randomly selected cases. The interobserver reproducibility of the extracted radiomic features was exceptionally high, yielding a median ICC of 0.999 (range: 0.492–1.000) across all features (Supplementary Fig. 2). Specifically, the median ICCs for the DWI and T2WI sequences were 0.999 (range: 0.658–1.000) and 0.999 (range: 0.492–1.000), respectively. To ensure the robustness of the constructed model, features exhibiting ICC ≤ 0.75 were systematically excluded prior to feature selection. Visual examples demonstrating the spatial agreement of ROIs between readers are provided in Supplementary Fig. 3.
Radiomic features
To evaluate the prognostic relevance of peritumoral microenvironments, concentric peritumoral zones were computationally generated by radially expanding tumor ROIs on GAN-enhanced T2-weighted imaging (T2WI) and diffusion-weighted imaging (DWI) sequences using a mask-based segmentation algorithm. The protocol employed 1 mm incremental expansions to create spatially resolved perilesional feature maps (Supplementary Fig. 4), with each layer independently analyzed for its predictive contribution.
Radiomic analysis encompassed both intratumoral and juxtatumoral regions, extracting morphometric (3D shape descriptors), intensity-based (histogram metrics), and high-order texture features. Texture characterization utilized computational operators including gray-level co-occurrence matrix (GLCM), gray-level run length matrix (GLRLM), gray-level size zone matrix (GLSZM), and neighborhood gray tone difference matrix (NGTDM). Feature quantification was performed using PyRadiomics v3.0.1, adhering to Image Biomarker Standardization Initiative (IBSI) guidelines. Subregional features were aggregated via pre-fusion to optimize discriminative capacity.
Feature selection involved univariate filtering (p < 0.05), collinearity removal (Pearson’s r > 0.9), minimum redundancy maximum relevance (mRMR) algorithm (prioritizing 32 features), and LASSO regression with 10-fold cross-validation (λ optimization) to retain statistically significant predictors (p < 0.05). For classification, machine learning models—Support Vector Machine (SVM) for linear tasks, Random Forest, ExtraTrees, and LightGBM for nonlinear structures—were optimized via grid search and 5-fold cross-validation. Both intratumoral (Intra) and peritumoral (PeriXmm, where “X” denotes radial distance) signatures were generated using identical pipelines.
Development of the 2.5D deep learning framework
A 2.5D representation was generated by cropping slices adjacent to the largest tumor cross-section along the superior-inferior axis. The selected slice range included ± 1, ± 2, and ± 4 slices around the central plane to provide rich spatial context. Dual-sequence inputs comprised DWI and T2WI as separate channels, supplemented by a synthetic third channel generated through Mixup fusion: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${c}{{mixup}}=\frac{{c}{{DWI}}+{c}_{T2{WI}}}{2}$$\end{document}cmixup=cDWI+cT2WI2. These tri-channel inputs were stacked for model training, with each 2D slice assigned patient-level annotations. Transfer learning leveraged pre-trained ImageNet weights (ResNet101/VGG19/DenseNet201) to enhance feature representation efficiency under limited data constraints (Supplementary 1A).
A multi-instance learning (MIL) paradigm aggregated slice-level predictions through prediction likelihood histograms (PLH) and TF-IDF-weighted Bag-of-Words (BoW), harmonizing deep learning outputs with conventional radiomic descriptors (Supplementary 1B). To ensure strict reproducibility and prevent data leakage, all feature selection steps, including the construction of TF-IDF dictionaries and PLH bin definitions, were derived exclusively from the training dataset. Post-aggregation, the MIL signature underwent feature refinement via t-tests, Pearson correlation filtering, and LASSO regularization performed on the training data. Classifiers (SVM/ExtraTrees/Random Forest) were optimized using 5-fold cross-validation and grid search hyperparameter tuning. Crucially, the synthetic minority over-sampling technique (SMOTE) was applied only to the training folds within each cross-validation loop to avoid contaminating the validation data. The ExtraTrees classifier was ultimately selected for the final model construction as it yielded the highest performance on the separate internal validation set. Ensemble validation further corroborated MIL robustness (Supplementary 1C).
Model evaluation integrated multivariate logistic regression (retaining features with p < 0.05 from univariate screening), combining clinical variables, peritumoral radiomics, and MIL outputs. Performance metrics included ROC-AUC analysis, Hosmer–Lemeshow calibration, and decision curve analysis (DCA) for clinical utility quantification. Final model selection prioritized validation cohort performance to ensure comparative fairness. The workflow schematic is shown in Fig. 2.

Fig. 2: Study workflow schematic
Analytical methodology
Statistical analyses utilized Python 3.7.12 (OnekeyAI platform). Categorical variables were compared using chi-square tests, and continuous variables via t-tests or Mann–Whitney U tests. Evaluation included ROC-AUC, calibration curves, and DCA. To construct the final Combined Model and avoid the curse of dimensionality, a stepwise integration strategy was adopted. First, independent clinical predictors were identified via univariate and multivariate logistic regression analyses restricted strictly to clinical variables. Only clinical features maintaining statistical significance (p < 0.05) in this multivariate analysis were retained. Subsequently, the identified clinical predictor (Tumor Length), the Radiomics Signature, and the MIL Signature were treated as independent biomarkers and entered into a final multivariate logistic regression analysis to generate the nomogram.