Section 2 of 8
Multi-omics integration and model construction
Jiao Meng, Wei Zhang, Hanmin Wang, Zihan He, and Zhuxian Zhang · about 15 minutes
Metabolomics biomarkers
Metabolomics, through the quantitative analysis of all small-molecule metabolites (typically <1500 Da) in biological fluids, directly reflects the final biochemical phenotype of an organism under specific physiological or pathological conditions. It serves as a bridge between genotype and clinical manifestation. The metabolite changes identified by metabolomics profoundly illustrate the complex pathophysiological states during the progression of DKD, such as glomerular hypertension, inflammation, oxidative stress, and mitochondrial dysfunction. Urine, as a non-invasive sample that can directly reflect the local metabolic state of the kidneys, is widely used in this field. Studies have shown that urine metabolite profiling can not only distinguish DKD from non-diabetic kidney disease, but also reveal key molecules associated with disease progression (He et al., 2025). For example, inositol has been identified as a novel prognostic biomarker significantly associated with DKD progression (Kwon et al., 2023).
Blood samples provide systemic metabolic information. Research has found that specific plasma metabolites—including lipids, amino acids, and carbohydrates—undergo significant changes at different stages of DKD, closely related to oxidative stress, inflammation, and fibrosis (Hasegawa and Inagi, 2021). A study using LC-MS/MS technology identified several differential metabolites, including lactate, L-ornithine, and L-tryptophan, and pointed out that disruptions in pathways such as amino acid biosynthesis and arginine-proline metabolism are among the core pathological mechanisms of DKD (Chen et al., 2025).
In addition, emerging metabolomics research on extracellular vesicles (EVs) reveals another dimension of biomarkers. By analyzing plasma EVs from DKD patients at different stages, researchers have discovered stage-specific metabolomic changes and identified several stable differential metabolites, including docosahexaenoic acid (DHA) and arachidonic acid (AA). This indicates that the metabolite composition within EVs can itself serve as a novel biomarker reflecting disease status (Pan et al., 2022).
In summary, metabolomics has uncovered groups of potential biomarkers from multiple levels—urine, blood, and EVs—that can warn of DKD risk and reveal its molecular mechanisms. However, the diagnostic performance of any single metabolomic biomarker remains limited by external confounders such as diet, drugs, and gut microbiota, underscoring the need for multi-omics integration.
Genomics-derived biomarkers
Genomics mainly seeks biomarkers from the perspectives of genetic susceptibility and epigenetic regulation, helping to identify populations at high risk for DKD and to explain disease heterogeneity. Genome-wide association studies (GWAS) have identified multiple genetic loci associated with the risk of developing DKD in patients with T2DM. Among these, the ELMO1 gene is the most extensively studied and has been validated as a DKD susceptibility gene across different ethnic groups. It is involved in key pathological processes of DKD such as apoptosis, inflammation, and fibrosis (Azarboo et al., 2024). In addition, genetic loci such as ACACB, CARS2, APOL1, and UNC13B are also thought to be associated with DKD susceptibility (Pan et al., 2023). Although these single nucleotide polymorphisms (SNPs) alone are insufficient for direct diagnosis of DKD, they have significant potential for assessing individual genetic susceptibility.
Epigenetic modifications act as a bridge between genetics and environment, offering new perspectives for understanding the mechanisms of DKD and for discovering new biomarkers. Major epigenetic modifications include DNA methylation, histone modification, and non-coding RNAs (ncRNA), which reversibly regulate gene expression and are easily influenced by environmental factors such as high glucose levels (Gu, 2019). ncRNAs do not code for proteins, but they regulate gene expression through various mechanisms and participate in the development and progression of DKD.
miRNAs are a class of short-chain ncRNAs that regulate gene expression post-transcriptionally by binding to target messenger RNAs. In DKD, the expression of various miRNAs changes significantly, participating in the regulation of core processes such as the TGF-β signaling pathway, epithelial-mesenchymal transition (EMT), and fibrosis (Sandholm et al., 2023). miRNAs are stably present in body fluids such as serum, plasma, and urine, making them well-suited as non-invasive “liquid biopsy” biomarkers (Uil et al., 2021). Research has identified multiple miRNAs associated with DKD; for example, miR-21 is regarded as an early marker of type 1 diabetic nephropathy (Fouad et al., 2020), and its promotion of fibrosis has been confirmed in animal models of T2DM-associated nephropathy (Dhas et al., 2023). miR-192-5p is significantly downregulated in renal tissue of DKD patients, and its overexpression can alleviate hyperglycemia-induced cellular damage (Wang et al., 2023). miR-145-5p is specifically overexpressed in DKD models, and its stability and detectability in urinary exosomes are favorable (Han et al., 2023). Other miRNAs such as miR-377, miR-29a, and miR-126 have also been explored for their value in early diagnosis (Yun et al., 2026; Mansouri et al., 2022).
lncRNAs are ncRNAs longer than 200 nucleotides and play critical roles in kidney diseases. In DKD, their aberrant expression can impact renal cell function by acting as miRNA sponges, recruiting modifying enzymes, and more (Ignarski et al., 2019). For example, lncRNA 153 is significantly downregulated in both renal tissue and serum of DKD patients, with its expression correlated with podocyte injury and disease severity (Yu et al., 2025). IFNG-AS1 is significantly upregulated in DKD patients and is associated with inflammatory responses and renal injury, making it a “pro-inflammatory biomarker.” In contrast, TH2LCRR is notably downregulated, serving as an ideal “protective biomarker.” (Hosseini et al., 2025a) lncRNAs SNHG1 and CRNDE participate in regulating the Th17/Treg cell balance, contributing to the identification of early DKD transition (Hosseini et al., 2025b). Serum lnc458 is elevated in DKD patients and positively correlates with the severity of proteinuria (Yang et al., 2025). These findings all suggest the potential of lncRNAs as diagnostic biomarkers, though further validation is needed for clinical translation.
circRNAs have a circular structure, exhibit high stability, are enriched in blood and urine, and participate in the DKD process by acting as “molecular sponges,” among other mechanisms (Benitez et al., 2024). For example, Circ-0000953 is significantly reduced in the renal tissue of both DKD patients and mouse models; its expression correlates with urinary microalbumin, serum creatinine, and eGFR, and changes can occur before obvious kidney injury symptoms develop, indicating potential for early screening (Liu et al., 2024a). In addition, circ_0003928 and circ_0068087 have also been reported to be involved in the onset and progression of DKD (Liu et al., 2022, 2024b).
In summary, from genetic susceptibility markers to environmentally regulated epigenetic markers (especially various ncRNAs), genomics provides a wealth of information and tools for the early identification, risk stratification, and mechanistic understanding of DKD. Nevertheless, genomic markers alone offer static risk stratification without temporal sensitivity, and epigenetic markers are context-dependent and reversible; thus, their integration with functional omics layers is essential for robust early detection.
Construction of composite diagnostic models
A single biomarker often fails to reflect complex DKD pathological processes, showing limited sensitivity and specificity. Integrating multi-omics data with machine learning to build composite diagnostic models is a promising approach. Fig. 1 illustrates the conceptual framework of multi-omics integration for early DKD diagnosis. At the biomarker level, five distinct omics layers each contribute unique, complementary diagnostic dimensions: genomics provides static risk stratification (e.g., ELMO1, APOL1, ACE I/D); epigenetics offers dynamic, reversible regulatory signatures (e.g., DNA methylation, histone modifications); transcriptomics captures post-transcriptional control networks (e.g., miR-192-5p, lncRNA-153, circ-0000953); proteomics reflects direct tissue injury evidence (e.g., KIM-1, NGAL); and metabolomics delivers real-time functional phenotypes (e.g., inositol, TMAO, DHA/AA in EVs). Nevertheless, each layer possesses inherent limitations when used in isolation—genomic markers are static and lack temporal sensitivity; epigenetic markers are context-dependent and reversible; transcriptomic markers lack direct phenotypic specificity; proteomic markers often emerge only after substantial structural damage; and metabolomic markers are susceptible to external confounders such as diet and drugs. Multi-omics integration overcomes these individual limitations by fusing their complementary contributions through feature selection, cross-layer validation, and pathway mapping, ultimately constructing predictive models that outperform single-omics approaches in sensitivity and specificity (Fig. 1).

Fig. 1: Integrated multi-omics framework for early diagnosis and risk stratification of diabetic kidney disease. This schematic illustrates the complementary contributions and inherent limitations of five omics layers in DKD. Genomics provides stable, germline-encoded risk predisposition (e.g., ELMO1, APOL1, ACE I/D); epigenetics captures reversible and context-dependent regulatory states (e.g., DNA methylation, histone modifications); transcriptomics reflects post-transcriptional control networks (e.g., miR-192-5p, lncRNA-153, circ-0000953); proteomics offers direct evidence of tissue injury (e.g., KIM-1, NGAL); and metabolomics delivers real-time functional readouts of systemic metabolism (e.g., inositol, TMAO, DHA/AA in EVs). In contrast to single-omics approaches, which suffer from limited sensitivity and specificity, this multi-omics integration employs feature selection, cross-layer validation, and pathway mapping to fuse heterogeneous data. Machine learning classifiers subsequently weight each layer by its incremental contribution to diagnostic accuracy. Within the current predictive model, the typical feature importance ranking was metabolomics > proteomics > epigenetics > transcriptomics > genomics, thereby enabling superior early DKD detection and risk stratification compared with any single-omics strategy.Abbreviations: DKD, diabetic kidney disease; EVs, extracellular vesicles; TMAO, trimethylamine N-oxide; DHA/AA, docosahexaenoic acid/arachidonic acid; KIM-1, kidney injury molecule-1; NGAL, neutrophil gelatinase-associated lipocalin; ACE I/D, angiotensin-converting enzyme insertion/deletion.
Multi-omics technologies can systematically integrate diverse information such as genomics and metabolomics, comprehensively analyzing disease mechanisms—such as identifying differentially expressed genes (like FN1, ALDH2) and associated pathway abnormalities (Lin et al., 2025; Liu et al., 2025; Luo et al., 2024).For instance, combining urinary metabolomics and peptidomics identified stepwise-regulated metabolites and peptides (e.g., UMOD, SERPINA1) that improve early DKD diagnosis (Jiang et al., 2023). Machine learning algorithms efficiently process high-dimensional data and select feature biomarkers, overcoming reliance on single indicators.
Recent studies utilized routine blood and biochemical parameters (e.g., TyG index, creatinine) to construct logistic regression models for early DKD screening (Yong et al., 2026), and integrated inflammatory–metabolic indices (e.g., UHR, NHR, SII) into XGBoost models with robust performance (Liu et al., 2026). Importantly, SHAP (Shapley Additive Explanations) analysis—based on cooperative game theory—quantitatively decomposes model predictions and assigns a contribution value to each biomarker, addressing the “black box” limitation of machine learning. As shown in Fig. 1, each omics layer is weighted by its incremental contribution to diagnostic accuracy, with a typical feature importance ranking of Metabolomics > Proteomics > Epigenetics > Transcriptomics > Genomics. For instance, SHAP analysis ranked eGFR, albumin, and C3 as top contributors to DKD prognosis, while peptidomic and metabolomic markers provide complementary, quantifiable information (Qian et al., 2026). Such interpretable frameworks enable precise quantification of each biomarker's diagnostic contribution, offering a strategy for translating multi-omics biomarkers into clinically actionable models.
Building upon the multi-omics integration framework illustrated in Fig. 1, where heterogeneous biomarker contributions are systematically fused through machine learning, we translated this conceptual model into a clinically actionable stratified diagnostic algorithm (Fig. 2). Based on the characteristics of the various biomarkers described above, we propose a stratified integrated diagnostic algorithm that systematically incorporates serum, urine, and genomic biomarkers according to the clinical application scenarios. This algorithm is centered on the four-dimensional complementarity of “genetic–structural–functional–mechanistic” markers and encompasses five progressive tiers: the risk prediction tier (Layer 0) incorporates SNPs (such as ELMO1 and APOL1) and epigenetic markers to assess individual DKD genetic susceptibility at the initial diagnosis of diabetes; the primary screening tier (Layer 1) combines urinary podocyte markers (Nephrin, Podocalyxin) with tubular markers (KIM-1, L-FABP, and NGAL) to identify structural injury even during normoalbuminuric stages; the diagnostic tier (Layer 2) introduces urinary metabolomic features (such as myo-inositol) and plasma amino acid profiles, supplemented by miRNA (miR-21 and miR-192-5p) and lncRNA (IFNG-AS1 and TH2LCRR) to aid in differentiating DKD from non-diabetic kidney disease; the prognostic monitoring tier (Layer 3) dynamically tracks urinary MCP-1/NGAL, plasma EVs metabolites (DHA and AA), and circRNA (circ-0000953) changes, in combination with eGFR slope and UACR trends to assess risk of progression; the therapeutic decision-making tier (Layer 4) uses serum metabolites (lactate, L-ornithine, and L-tryptophan) to monitor recovery of amino acid pathways, guiding efficacy evaluation and regimen adjustment for drugs such as RAS inhibitors and SGLT2i. This stratified framework not only addresses the lack of specificity inherent in single biomarkers (see Table 1 for details) but also clarifies the complementary roles of multi-omics biomarkers at different clinical nodes.

Fig. 2: Multi-omics-based clinical workflow and interpretable risk stratification for diabetic kidney disease. This schematic presents a stepwise diagnostic and prognostic algorithm (Layer 1: Screening to Layer 5: Model Interpretability) integrating genomic, epigenomic, metabolomic, and proteomic biomarkers from serum, urine, and tissue specimens. Layer 1 (Screening) identifies initial genomic risk via SNPs (e.g., ELMO1, APOL1) and urinary damage markers (NGAL, L-FABP); abnormal results trigger progression to Layer 2 (Diagnosis), which confirms DKD via urinary/plasma miRNAs, circRNAs, and lncRNAs. Confirmed cases proceed to Layer 3 (Prognosis), wherein serial monitoring of plasma/EVs markers (MCP-1, miR-21, DHA/AA) defines quarterly (q3mo) follow-up. Layer 4 (Treatment) evaluates metabolic recovery through amino acid pathway restoration (lactate, L-ornithine, L-tryptophan) to guide therapeutic decisions. Layer 5 employs two complementary interpretability tools: (a) a global feature importance bar plot ranking omics-layer contributions (metabolomics > proteomics > epigenetics > genomics) to the predictive model; and (b) a SHAP force plot providing patient-level explanation, where deviations in key biomarkers (↑miR-21, ↑NGAL, ↓inositol, ↓eGFR) drive individual risk classification toward “High-risk DKD.”Abbreviations: DKD, diabetic kidney disease; EVs, extracellular vesicles; L-FABP, liver-type fatty acid-binding protein; MCP-1, monocyte chemoattractant protein-1; miRNA, microRNA; NGAL, neutrophil gelatinase-associated lipocalin; SHAP, SHapley Additive exPlanations; SNP, single nucleotide polymorphism.Symbol and Line Legend: Sample icons: Yellow = urine; red = plasma/serum; pink = tissue biopsy. Solid arrows: Defined workflow progression; Dashed arrows:Alternative workup or repeat monitoring pathways. Colored arrows: Red (↑) denotes biomarkers whose elevated levels push prediction toward high-risk; Blue (↓) denotes markers whose reduced levels are associated with adverse renal outcomes (e.g., ↓inositol, ↓eGFR). Circular nodes & timeline: Key metabolites in the amino acid recovery pathway; quarterly (3-month) monitoring intervals for dynamic markers.
The construction of such models usually follows a core technical process that mirrors the integration pipeline depicted in Fig. 1. First, multi-omics data and traditional clinical indicators (such as UACR, eGFR) are collected from samples like plasma and urine, and undergo rigorous quality control and standardization to ensure comparability. Second, core biomarker panels are identified using algorithms such as LASSO regression and random forest, which correspond to the feature selection stage of Fig. 1. Third, SHAP analysis is applied to quantify the contribution of each omics biomarker to model prediction, thereby identifying key global drivers (e.g., metabolomic biomarkers often contribute most) as well as localized individual explanations (e.g., the interaction effect between miR-21 and myo-inositol in specific patients), thus revealing synergistic and redundant relationships among biomarkers—consistent with the cross-layer validation and pathway mapping stages of Fig. 1. Next, these multi-omics biomarkers are integrated with traditional metrics to build predictive models. Finally, the performance indices of the models, such as sensitivity, specificity, and area under the curve (AUC), are evaluated in independent cohorts and optimized through methods such as cross-validation to ensure generalizability.
Ultimately, these biomarkers are integrated with traditional clinical indicators (UACR, eGFR) to construct interpretable stratified prediction models, with generalizability further optimized via cross-validation and external validation. Existing research has demonstrated the effectiveness of this approach—for example, by integrating multi-omics data from the GEO database, key biomarkers such as AGR2 and CCR2 have been identified using WGCNA and LASSO regression, with external validation confirming satisfactory diagnostic performance (Wang et al., 2024; Liu et al., 2023). Future research could further incorporate SHAP analysis, not only improving the predictive accuracy of the model but also enhancing its clinical interpretability, thus transforming the “black box” model into a decision support tool that physicians can understand.
In summary, multi-omics integration can characterize disease heterogeneity at the molecular level, overcoming the limitations of traditional staging, while machine learning enhances the objectivity of biomarker selection and the generalizability of models. The combination of the two not only shows promise for addressing clinical challenges such as delayed early diagnosis and ambiguous subtyping of DKD, but also offers new avenues for precision medicine. Future research should focus on intelligently integrating information from multiple sources to build scenario-based precision diagnostic pathways.