Work overview

Section 04 of 05

Discussion

When Literature Priors Are Removed: Clinical Rank Shifts and Transancestral Stability in Antipsychotic Low-Density Lipoprotein (LDL)-Risk Modelling

Metadata pending adapter verification · 2026

Contents

Section 04 of 05

  1. 01Introduction
  2. 02Materials and methods
  3. 03Results
  4. 04Discussion
  5. 05Conclusions
Text size
Work overview

Section 4 of 5

Discussion

Metadata pending adapter verification · about 9 minutes

Transancestral robustness

The five ancestry runs shared the same Ki matrix, DDD values, receptor-to-gene map, drug list, and literature-derived prior dictionary. Accordingly, consistency across runs demonstrates structural reproducibility of model behavior, not independent biological replication or ancestry-specific clinical LDL risk. The main finding of this study is that the LDL-related drug ranking was highly stable across ancestry-defined datasets, while the receptor-level explanation of the score was strongly sensitive to the weighting scheme. The weighted and uniform models produced mean cross-ancestry rank correlations of 0.996 and 0.995, respectively. The weighted top-10 set was identical across all ancestry pairs, whereas the uniform model produced a lower mean top-10 Jaccard value of 0.758. This pattern suggests that ancestry-specific LDL TWAS inputs caused relatively small perturbations to the broad drug ordering but more visible changes when the model was allowed to rely on equal receptor priors.

The gene-level results were even more consistent. All 39 genes represented in the exported direction-consistency analysis changed contribution share in the same direction across all five datasets. DRD2, DRD3, and DRD4 consistently gained relative share, while HRH1, HTR2A, and HTR2C consistently lost share. DRD2 was the leading uniform-model contributor in every ancestry. These findings support a stable direction of model sensitivity rather than an ancestry-specific effect driven by a single dataset.

The result is notable in the context of multi-ancestry lipid genetics, where ancestry differences in allele frequencies, linkage disequilibrium, prediction-model performance, and fine-mapping resolution can affect gene-level association patterns [14-16]. However, the interpretation must remain narrow. The five analyses shared the same drug list, Ki values, defined daily doses, receptor mappings, contribution formula, and prior dictionary. The perfect directional consistency therefore reflects both the ancestry-specific TWAS inputs and a common pharmacological structure. It demonstrates transancestral stability of model behavior, not transancestral validation of causal receptor mechanisms.

The appropriate conclusion is that ancestry-specific LDL TWAS inputs produced only modest changes in relative drug ranking within this framework. This supports evaluation of a common reference ranking for research communication, but it does not establish that ancestry-specific clinical adjustment is unnecessary. Clinical generalizability requires patient-level outcome validation in diverse populations.

Clinical context and prior dependence

The literature-weighted model had greater qualitative concordance with known metabolic hierarchies than the uniform model, but its agreement with clinical evidence was qualitative rather than exact. Clozapine remained the highest-ranked drug across all ancestry datasets, and olanzapine remained in the upper tier. These findings are compatible with clinical evidence showing greater metabolic liability for clozapine and olanzapine, including greater average adverse effects on weight, glucose, and selected lipid outcomes than for several newer agents [1,2,4]. The model also placed aripiprazole lower than many of the drugs with stronger histaminergic and serotonergic contributions.

The weighted model nevertheless ranked chlorpromazine second and asenapine third in the cross-ancestry consensus ordering, while ziprasidone occupied a mean rank of 5.80 in the broader weighted summary. These placements should not be presented as a validated clinical hierarchy. They show that the model retains a mixture of pharmacological affinity, defined daily dose, TWAS evidence, nominal significance bonuses, and normalization choices. External clinical studies provide context for the output but do not validate the exact order of all 52 drugs.

Removing the literature priors made this distinction clearer. The uniform model moved ziprasidone, aripiprazole, and haloperidol toward numerically smaller ranks, which correspond to higher predicted risk because rank 1 denotes the highest predicted risk. It also moved olanzapine and quetiapine toward numerically larger ranks. These movements are discordant with the usual clinical interpretation of these drugs and demonstrate that a uniform model should not be used alone for treatment selection or patient counselling.

The class-level reallocation helps explain the behavior. The weighted model placed substantial relative emphasis on histaminergic and serotonergic systems, especially HRH1, HTR2C, and HTR2A. The clinical relevance of H1 and 5-HT2C mechanisms is supported by receptor-binding and metabolic literature [9-12]. The network meta-analysis [2] also found that H1, M1, and M3 receptor occupancies were associated with greater glyco-metabolic deterioration. The weighted model therefore embeds prior knowledge that is broadly consistent with known antipsychotic metabolic patterns. This makes it more interpretable, but also partly circular.

Dual-reporting recommendation

The results support reporting both models rather than selecting one model and treating it as definitive. We recommend dual reporting: the literature-weighted model as the clinically informed reference and the uniform model as a sensitivity analysis of receptor-prior assumptions. The literature-weighted model should be used as the primary clinically informed reference because it incorporates established pharmacological knowledge and better preserves the expected prominence of agents with substantial metabolic liability. It may be useful for comparative research communication and for organizing hypotheses about drug-level metabolic burden.

The weighted model should not be described as clinically validated. Its receptor weights were manually specified from prior literature, its dose term is a defined daily dose proxy rather than a measure of individual exposure, and its output has not been calibrated against longitudinal LDL, weight, glucose, diabetes, metabolic syndrome, or cardiovascular outcomes. Clinical guidance supports monitoring and management of metabolic risk during antipsychotic treatment, but it does not establish this score as a clinical decision instrument [6-8].

The uniform model should be retained as a sensitivity and discovery-oriented analysis. It asks how much the drug ranking and receptor attribution depend on the literature priors. It can reveal genes and pathways that are suppressed by prior weighting, including DRD2, DRD3, DRD4, HTR2B, and ADRB1. It should not be described as unbiased, because equal weighting is itself a prior assumption. It also restores contribution from genes absent from the weighted dictionary, so the comparison is a change in both the strength and the scope of the receptor prior.

A dual-reporting framework is therefore more informative than a single model. The weighted model communicates the consequences of incorporating established pharmacological evidence. The uniform model reveals how the conclusions change when that evidence is intentionally withheld. Agreement between the models supports robustness. Disagreement identifies areas where mechanistic interpretation is prior-dependent.

HRH1, DRD2, and ADRB1 as hypotheses

HRH1 was the strongest ranking-sensitive component in the leave-one-gene-out analysis. This result is compatible with prior evidence linking H1 receptor affinity to weight gain and with clinical pharmacology studies associating H1 occupancy with metabolic deterioration [2,9]. The present analysis extends that observation by showing how removing HRH1 changes the relative positions of particular drugs within the reconstructed LDL ranking. It does not show that HRH1 is the sole or dominant causal mediator of LDL change.

DRD2 became the leading gene-level reconstructed contribution under uniform weighting in every ancestry. This finding is biologically plausible because dopaminergic pathways influence reward, feeding behavior, endocrine signaling, and metabolic regulation. Nevertheless, the result is generated by the interaction of Ki values, defined daily dose, TWAS evidence, p-value scaling, and the removal of differential priors. It should therefore be treated as a testable hypothesis rather than evidence that D2-related mechanisms are clinically more important than histaminergic or serotonergic mechanisms.

ADRB1 was the strongest high-TWAS/low-weight exploratory, single-ancestry candidate, with an absolute z-score of 6.079 in the European dataset and a literature weight of 0.08. Adrenergic signaling has established roles in adipose-tissue metabolism and energy expenditure, and beta-1 adrenergic receptors may contribute to human brown-adipocyte metabolic activity [24,25]. These studies provide biological plausibility but do not demonstrate that ADRB1 mediates antipsychotic-associated LDL changes. The present ADRB1 signal was supported in only one ancestry, so it does not yet meet a multi-ancestry replication standard. Independent TWAS replication, colocalization, tissue-specific analysis, fine-mapping, and functional testing are needed before assigning clinical or causal importance.

Relation to the broader series

The LDL analysis complements the earlier comparison of four TWAS transformations and the receptor-prior reallocation analysis. The scaling-method study showed that the choice of TWAS transformation mainly changed score magnitude and had a smaller effect on drug ordering [19]. The prior-removal analysis showed that rank stability can coexist with substantial changes in receptor-level explanation and identified ADRB1 as an underweighted candidate [20]. The HDL Stage 3 analysis examined whether these effects were stable across ancestry-defined datasets and found a similar distinction between stable model behavior and clinically discordant individual rankings [21].

The LDL results strengthen the generalizability of that conceptual framework across lipid traits. The recurring pattern is not that every receptor finding is identical between LDL and HDL, but that removing priors consistently exposes a difference between ranking robustness and mechanistic robustness. That distinction is relevant to other pharmacological models that combine empirical drug features with literature-derived biological weights.

Limitations

Several limitations should be considered. Most importantly, this score is a computational prioritization index rather than a validated clinical risk predictor. No patient-level LDL changes, triglyceride changes, weight trajectories, glucose outcomes, incident diabetes, metabolic syndrome diagnoses, or cardiovascular events were used for validation.

Second, defined daily dose is a standardized dose proxy. It does not represent the prescribed dose for an individual patient, treatment adherence, treatment duration, plasma exposure, receptor occupancy, or tissue concentration. The Ki-DDD term is therefore useful for comparative modeling but cannot be interpreted as a direct estimate of pharmacological exposure.

Third, the pipeline used minimum Ki aggregation. Selecting the lowest reported Ki value may be useful for identifying potential binding interactions, but it can overemphasize an isolated measurement and may not represent the central affinity of a drug across assays. A median, mean, or geometric-mean sensitivity analysis was not included in the saved Stage 3 run. Alternative aggregation strategies, including median, mean, or hierarchical models, should be evaluated.

The internal nature of the Ki database limits record-level reproducibility. In addition, the literature-derived receptor-weight dictionary is not independent of the pharmacological evidence used to interpret metabolic liability; its use therefore introduces a form of circularity into the weighted model. Conversely, the uniform model is not assumption-free because equal weighting is itself a prior.

Fourth, fuzzy ligand matching and generic receptor labels introduce uncertainty. A generic alpha-1 label was mapped to ADRA1A, a generic alpha-2 label to ADRA2A, and a generic muscarinic label to CHRM1. These operational choices may not accurately represent subtype-specific binding. The final workflow should preserve alternative matches and manual adjudication status.

Fifth, missing TWAS values were imputed. The upstream runs contained 97 to 175 nominally significant rows and approximately 210 to 239 imputed TWAS rows per ancestry. The reconstruction also reapplied mean negative-log10 p-value imputation for missing scored genes. This procedure avoids automatically penalizing drugs with incomplete coverage, but it may reduce genuine heterogeneity and introduce common contribution into rows without directly observed TWAS values. The saved Stage 3 outputs did not include a separate analysis excluding imputed rows, so the magnitude of this influence could not be quantified.

Sixth, the TWAS transformations used p-values and absolute z-scores. They therefore ignored the direction of association. The score identifies genes with strong statistical evidence but cannot determine whether genetically predicted expression is associated with higher LDL, lower LDL, or a clinically adverse direction. The term “LDL-associated contributor” is more appropriate than “LDL risk gene.”

Seventh, the nominal p-value threshold of .05 was not corrected for multiple testing. The high-TWAS/low-weight screen was exploratory and was not equivalent to genome-wide significance, false-discovery-rate control, colocalization, or causal evidence.

Eighth, the five ancestry analyses were not fully independent replications of the complete model because they shared the same Ki data, defined daily dose values, receptor mappings, drug list, and prior dictionary. The strong consistency of the rank changes therefore provides evidence of structural reproducibility but cannot by itself establish that the same causal mechanisms operate in every ancestry.

Ninth, the leave-one-gene-out calculation used reconstructed receptor contributions and did not fully reproduce all drug-level multipliers in the saved final score. It should therefore be described as sensitivity of the reconstructed contribution-based ranking. Future analyses should recompute the complete final score after gene removal.

The saved Stage 3 outputs did not preserve complete package-version metadata, which limits exact reproduction of the computational environment. The negative-log10 p-value cap of 25 was used as a numerical safeguard, but its sensitivity was not formally tested in the saved Stage 3 run.

Finally, the present Stage 3 output did not contain an independent patient-level clinical comparator for LDL. Clinical alignment was therefore qualitative. A future study should prespecify a comparator based on independent clinical data and evaluate discrimination, calibration, rank agreement, and outcome-specific performance.