Work overview

Section 02 of 06

Materials and methods

Serum metabolic signatures associated with maximal voluntary ventilation in native high-altitude Tibetans

Tao Zhou, Jiawei Yang, Qiong Zhang, Haichen Zhang, Lening Chen, Shusheng Luo, Qianqian Xiao, Qinghe Meng, Jianjun Jiang, Labasangzhu Labasangzhu, Dunyou Dunyou, Danbalangjie Danbalangjie, Weidong Hao, and Xuetao Wei · 2026

Contents

Section 02 of 06

  1. 01Introduction
  2. 02Materials and methods
  3. 03Result
  4. 04Discussion
  5. 05Conclusion
  6. 06Supplementary Information
Text size
Work overview

Section 2 of 6

Materials and methods

Tao Zhou, Jiawei Yang, Qiong Zhang, Haichen Zhang, Lening Chen, Shusheng Luo, Qianqian Xiao, Qinghe Meng, Jianjun Jiang, Labasangzhu Labasangzhu, Dunyou Dunyou, Danbalangjie Danbalangjie, Weidong Hao, and Xuetao Wei · about 12 minutes

Study population and sample screening

A total of 202 native high-altitude Tibetan adults were initially enrolled in the present study. Strict sample quality control and exclusion criteria were applied to guarantee data reliability. Participants with missing maximal voluntary ventilation (MVV) data (n = 8), a documented history of pulmonary diseases (n = 12), or incomplete core clinical covariates (n = 14) were excluded. After systematic screening, 168 participants with fully complete clinical, phenotypic, and metabolomic data were included in the final analytical cohort, with no missing values observed in any variable. This investigation strictly adhered to the principles of the Declaration of Helsinki, fully safeguarding the life, health, dignity, and autonomy of all participants. Informed consent was obtained from each subject prior to their inclusion in the research. This study was approved by the Medical Ethics Review Committee of Tibet University (2023SQ001).

Clinical covariate selection and standardized integration

Based on well-established confounding factors affecting pulmonary function in high-altitude populations, a total of 20 preliminary clinical covariates covering six domains were recruited for confounder adjustment, including demographics, anthropometric parameters, socioeconomic status, lifestyle factors, sleep quality, and medical history. The full covariate list comprised: (1) demographics: age, integrated gender-menopause status; (2) anthropometric measurements: height, weight, body mass index (BMI), systolic blood pressure (SBP), diastolic blood pressure (DBP); socioeconomic indicator: per capita monthly income; (4) lifestyle factors: smoking status, cumulative smoking pack-years, drinking status, and metabolic equivalent of task; (5) sleep-related indicators: Breathing difficulty, Sleep latency or Sleep maintenance, Cough snoring, Subjective sleep quality, and Sleep medication; (6) chronic disease history: hypertension and dyslipidemia.

To reduce multicollinearity, eliminate data redundancy, and improve model interpretability, several covariates were rationally integrated and standardized according to physiological logic and epidemiological conventions. First, gender and menopausal status were combined into a unified reproductive characteristic variable to reflect sex-specific physiological differences in ventilatory function. Second, duplicate SBP and DBP measurements were averaged to obtain stable blood pressure values. Third, cumulative smoking pack-years were calculated from daily cigarette consumption and smoking duration to quantify long-term smoking exposure. Fourth, MET values were integrated from the frequency and duration of weekly light, moderate, and heavy physical activities to generate a continuous quantitative index of physical activity level.

Blood collection and serum metabolite profiling

Blood Collection and Serum Preparation

Fasting venous blood was collected from all participants using vacuum blood collection tubes. The collected blood was allowed to clot naturally at room temperature (~ 25 °C) for 30 min, followed by centrifugation at 4 °C and 3,000 rpm for 15 min. The supernatant (serum) was carefully aspirated, aliquoted into sterile centrifuge tubes, and stored at − 80 °C until extraction.

(2)Metabolite Extraction

Serum samples were thawed on an ice-water mixture. For each sample, 90 µL of serum was transferred to a 1.5 mL Eppendorf tube, and 360 µL of pre-chilled protein precipitant (methanol: acetonitrile = 2:1, v/v) containing a mixed internal standard (4 µg/mL) was added. Four stable isotope-labeled internal standards with confirmed vendor information and high purity were used for analytical quality control: L-2-chlorophenylalanine (C2001, Hengchuang Bio, Shanghai, 98.0% purity), succinic acid-d4 (293075-1G, Sigma, 98.0% purity), L-valine-d8 (HY-11124, Haoyuan Bio, Shanghai, 98.0% purity), and cholic acid-D4 (S22155-50 mg, Yuanye Bio, Shanghai, 98.0% purity) The mixture was vortexed vigorously for 1 min, ultrasonicated in an ice-water bath for 10 min, and incubated at − 40 °C overnight to ensure complete protein precipitation. Following incubation, samples were centrifuged at 12,000 rpm for 10 min at 4 °C. A 150 µL aliquot of the supernatant was filtered through a 0.22 μm organic-phase syringe filter, transferred to an LC autosampler vial, and stored at − 80 °C until LC-MS analysis. All samples were extracted on the same day to minimize batch effects. A quality control (QC) sample was prepared by pooling equal volumes of each extracted sample. A pooled quality control (QC) sample was prepared by mixing equal-volume supernatants extracted from all individual study samples to monitor analytical stability throughout the run.

(3)Liquid Chromatography–Mass Spectrometry (LC-MS) Conditions

Untargeted metabolomic profiling was performed using a Waters ACQUITY UPLC I-Class Plus ultra-high-performance liquid chromatography system coupled to a Thermo Q Exactive HF high-resolution mass spectrometer (UHPLC-HRMS). Chromatographic separation was achieved on an ACQUITY UPLC HSS T3 column (100 mm × 2.1 mm, 1.8 μm; Waters Corporation) maintained at 45 °C. The mobile phase consisted of (A) water with 0.1% formic acid and (B) acetonitrile, delivered at a fixed flow rate of 0.35 mL/min. The fixed injection volume was 3 µL. The injection volume was 3 µL. The gradient elution program was as follows: 0–2 min, 5% B; 2–4 min, 5%–30% B; 4–8 min, 30%–50% B; 8–10 min, 50%–80% B; 10–14 min, 80%–100% B; 14–15 min, 100% B; 15.0–15.1 min, 100%–5% B; 15.1–16 min, 5% B.

Mass spectrometric detection was performed in both positive and negative electrospray ionization (ESI) modes to maximize metabolite coverage. The key MS parameters were as follows: spray voltage, + 3,800 V (positive mode)/−3,000 V (negative mode); capillary temperature, 320 °C; auxiliary gas heater temperature, 350 °C; sheath gas flow rate, 35 arbitrary units (Arb); auxiliary gas flow rate, 8 Arb; S-lens RF level, 50. Full-scan data were acquired over the mass range m/z 70–1,050 at a resolution of 60,000 (full width at half maximum). Data-dependent acquisition (DDA) was employed for MS/MS fragmentation at a resolution of 15,000, with normalized collision energy (NCE) set to 10, 20, and 40 (stepped NCE).

(4)Quality Assurance and Quality Control (QA/QC)

A pooled QC sample was injected at the beginning and end of each analytical batch, and additionally after every 15 study samples, yielding a total of 20 QC injections across the run. The QC samples were used to monitor instrument stability, assess data reproducibility, and perform signal correction during data processing. Instrument stability was evaluated by calculating the Pearson correlation between QC samples; higher correlation coefficients indicated greater analytical stability. Prior to each analytical run, the mass spectrometer was calibrated using sodium formate to ensure mass accuracy. Features with a coefficient of variation (CV) > 30% across repeated QC injections were excluded to filter low-reproducibility metabolic features; this 30% CV threshold is a well-accepted standard for untargeted serum metabolomics. All internal standards were consistently detected with low CV values (5%–15%) across QC replicates, confirming reliable instrument stability. No gradient-concentration spiked QC samples were prepared, consistent with standard untargeted metabolomics discovery workflows.

(5)Data Processing and Metabolite Identification

Raw LC-MS data were processed using Progenesis QI version 3.0 (Nonlinear Dynamics, Newcastle, UK). The complete standardized data processing pipeline was strictly implemented as follows: baseline filtering, automatic peak detection, peak integration, software-default retention time alignment, peak grouping, and total ion current (TIC) normalization. For feature filtering, ion features with > 50% missing values across all samples were removed; this lenient threshold was adopted to retain low-abundance trace metabolites, dietary metabolites, and xenobiotics common in serum samples, while eliminating severely undetectable features. Remaining missing values were imputed using the half-minimum (Min/2) value of each feature, a standard method for signals below the limit of detection.

Metabolite identification was performed by matching accurate mass and MS/MS fragmentation spectra against four public and in-house databases: the Human Metabolome Database (HMDB), Lipidmaps, LuMet-Animal (an OE-specific animal metabolome database), and METLIN. A tiered mass tolerance strategy was applied to balance annotation confidence and coverage: precursor mass tolerance of ± 5 ppm and product ion tolerance of ± 10 ppm for rigorously curated HMDB and Lipidmaps databases to reduce false positives; precursor mass tolerance of ± 10 ppm and product ion tolerance of ± 20 ppm for LuMet-Animal and METLIN databases to maximize metabolome coverage.

The Progenesis QI scoring algorithm assigned a composite identification score (maximum 80 points) integrating four orthogonal dimensions: isotopic distribution matching (20 points), fragment ion matching (20 points), retention time consistency (20 points), and adduct ion plausibility (20 points). A composite score cutoff of ≥ 36 (45% of the total maximum score) was applied to retain reliably annotated features, ensuring sufficient matching evidence from at least two or three orthogonal identification dimensions.

A custom four-tier metabolite annotation confidence system (renamed Confidence Tier 1–4 to avoid confusion with community-standard MSI Level classification) was defined explicitly: Tier 1, retention time (RT) deviation ≤ ± 0.3 min AND Fragmentation Score ≥ 45; Tier 2, RT deviation ≤ ± 0.3 min AND 0 ≤ Fragmentation Score < 45; Tier 3, Fragmentation Score ≥ 45 without valid RT matching; Tier 4, Fragmentation Score < 45 without valid RT matching. Notably, this tier labeling only appears in supplementary annotation tables and is not mentioned in the main manuscript text. For key metabolites of interest, manual inspection of matching information was performed to further verify identification confidence. This untargeted discovery study did not purchase full-spectrum authentic reference standards for all metabolite validation; available reference standards were detected in neat solvent for RT and spectral matching rather than biological matrix spiking. A complete list of all identified metabolites with annotation-level information (database identifiers, mass error, retention time, identification score, and MS/MS fragment matching) is provided in Supplementary File 4, with fully revised and clearly defined table headers and detailed column explanations in supplementary readme files.

Data preprocessing and covariate collinearity diagnosis

A total of 5310 metabolic features were initially detected after peak picking. After missing value filtering and QC-CV filtering, 3381 high-quality annotated metabolic features were retained for all subsequent analyses. Prior to regression modeling, outcome (MVV) and all serum metabolite features were Z‑score standardized. Categorical clinical covariates were not standardized; only continuous clinical covariates underwent Z‑score transformation. Confounder adjustment was implemented via residualization: clinical covariates were regressed out from both MVV and individual metabolites, and the resulting residuals were used for subsequent univariate screening. For preliminary covariates, multicollinearity was assessed using the variance inflation factor (VIF). Variables with VIF < 10 were retained. After collinearity filtering, a total of 17 independent clinical covariates were finalized for subsequent model adjustment.

Stratified dataset partitioning

To avoid data leakage and ensure rigorous model validation, the entire cohort was stratified by MVV quartiles and randomly split into a training set (70%, n = 119) and an independent external validation set (30%, n = 49). All metabolite screening, variable selection, model training, and internal validation procedures were performed exclusively in the training set. The independent validation set was not involved in any screening or training processes and was only used to evaluate external model generalizability.

Multi-stage metabolite screening strategy

A multi-layered robust screening framework was constructed on the training set to screen serum metabolites associated with MVV. Prior to metabolite screening, we established a baseline linear regression model with MVV as the dependent variable and all predefined covariates as independent predictors to calculate covariate-adjusted MVV residuals; this baseline covariate-only model explained 22.85% of MVV variation (R2 = 0.2285). All categorical covariates were converted into dummy variables before regression fitting for standardized confounding elimination as shown in Table 1. In the first screening stage, univariate linear regression was performed individually for each metabolite against the adjusted MVV residuals, followed by three-tier hierarchical threshold filtering to preliminarily screen candidate metabolites. The R2 values of all single-metabolite univariate models ranged from 0.0000 to 0.1113, with a median of 0.0040 and a mean of 0.0081. Since no metabolite reached significance after FDR correction, we adopted successive nominal thresholds of (P < 0.05) and (P < 0.01) to expand the candidate pool for subsequent regularized regression. In the second stage, elastic net regularized regression (alpha = 0.5), lambda.1se) was adopted for high-dimensional dimensionality reduction and stable feature selection. To prevent null screening outcomes and guarantee analytical robustness, a fallback screening rule was predefined: elastic net modeling was executed when no fewer than two candidate metabolites were available. If all metabolites obtained zero regression coefficients after regularization, the top five metabolites with the smallest univariate P values would be retained as alternative candidate features.

Variable | Training set | Independent validation set
SBP | 120.00(106.75, 135.50) | 128.00(113.00, 137.50)
DBP | 81.00(75.00, 88.75) | 87.00(79.50, 91.00)
Age | 58.00(41.50, 69.50) | 62.00(49.00, 70.00)
Household income |  | 
< 3000 (1) | 40(33.6%) | 13(26.5%)
3000–4999 (2) | 37(31.1%) | 16(32.7%)
5000–7999 | 21(17.6%) | 10(20.4%)
8000–9999(4) | 8(6.7%) | 2(4.1%)
10,000–14,999(5) | 10(8.4%) | 6(12.2%)
> 15,000(6) | 3(2.5%) | 2(4.1%)
MET | 1070.00 (626.50, 3066.00) | 1197.00 (579.00, 2856.00)
Sleep latency |  | 
None(1) | 76(63.9%) | 29(59.2%)
Less than once a week(2) | 18(15.1%) | 9(18.4%)
Once or twice a week | 12(10.1%) | 6(12.2%)
Three or more times a week(4) | 13(10.9%) | 5(10.2%)
Sleep maintenance |  | 
None(1) | 70(58.8%) | 26(53.1%)
Less than once a week(2) | 28(23.5%) | 12(24.5%)
Once or twice a week | 9(7.6%) | 7(14.3%)
Three or more times a week(4) | 12(10.1%) | 4(8.2%)
Breathing difficulty |  | 
None(1) | 107(89.9%) | 39(79.6%)
Less than once a week(2) | 5(4.2%) | 6(12.2%)
Once or twice a week | 4(3.4%) | 1(2.0%)
Three or more times a week(4) | 3(2.5%) | 3(6.1%)
Cough snoring |  | 
None(1) | 89(74.8%) | 34(69.4%)
Less than once a week(2) | 18(15.1%) | 6(12.2%)
Once or twice a week | 7(5.9%) | 7(14.3%)
Three or more times a week(4) | 5(4.2%) | 2(4.1%)
Subjective sleep quality |  | 
Very good(1) | 43(36.1%) | 18(36.7%)
Fairly good(2) | 63(52.9%) | 22(44.9%)
Fairly bad | 10(8.4%) | 7(14.3%)
Very bad(4) | 3(2.5%) | 2(4.1%)
Sleep medication |  | 
None(1) | 115(96.6%) | 46(93.9%)
Less than once a week(2) | 2(1.7%) | 2(4.1%)
Three or more times a week(4) | 2(1.7%) | 1(2.0%)
Smoking status |  | 
Never smoker(0) | 86(72.3%) | 39(79.6%)
Ever smoker(1) | 33(27.7%) | 10(20.4%)
Pack_years | 0.00(0.00, 0.00) | 0.00(0.00, 0.00)
Alcohol drinking |  | 
Ever drinker(1) | 54(45.4%) | 20(40.8%)
Never drinker(2) | 65(54.6%) | 29(59.2%)
Hypertension |  | 
Yes(1) | 28(23.5%) | 9(18.4%)
No(2) | 91(76.5%) | 40(81.6%)
Dyslipidemia |  | 
Yes(1) | 13(10.9%) | 10(20.4%)
No(2) | 106(89.1%) | 39(79.6%)
Sex_menopause status |  | 
Male(0) | 45(37.8%) | 17(34.7%)
Premenopausal female(1) | 41(34.5%) | 14(28.6%)
Postmenopausal female(2) | 33(27.7%) | 18(36.7%)
MVV | 36.44(29.02, 50.72) | 40.03(30.75, 49.19)

Adjusted multivariable linear regression model

Candidate metabolites obtained from multi‑stage screening were incorporated into multivariable linear regression models, with Z‑score‑standardized MVV as the dependent variable and the 17 filtered clinical confounders fully adjusted. Metabolites with P < 0.05 in the fully adjusted model were considered independently associated with MVV. After multivariable adjustment, three serum metabolites were identified as core independent metabolic correlates of pulmonary ventilatory function. A sensitivity analysis was further performed by excluding the exogenous drug‑related metabolite (Isosorbide Dinitrate), repeating the full multi‑stage analytical pipeline to test the robustness of endogenous metabolite‑MVV associations.

Model validation and stability evaluation

Multiple validation approaches were adopted to comprehensively assess model performance and screening stability. Ten-fold cross-validation was conducted in the training set to evaluate internal predictive performance and calculate the average cross-validated R2. Bootstrap resampling analysis was further performed to quantify the stability and repeatability of metabolite screening outcomes. Finally, the optimized metabolite model was tested in the independent validation set to evaluate external generalization capability, and the corresponding predictive R2 was determined.