Work overview

Section 02 of 05

Materials and methods

When Literature Priors Are Removed: Clinical Rank Shifts and Transancestral Stability in Antipsychotic Low-Density Lipoprotein (LDL)-Risk Modelling

Ngo Cheung, Hoi-Ki Cheung, Yee-Wah Yu, and Yolanda Yuen-Ching Tsang · 2026

Contents

Section 02 of 05

  1. 01Introduction
  2. 02Materials and methods
  3. 03Results
  4. 04Discussion
  5. 05Conclusions
Text size
Work overview

Section 2 of 5

Materials and methods

Ngo Cheung, Hoi-Ki Cheung, Yee-Wah Yu, and Yolanda Yuen-Ching Tsang · about 8 minutes

Data sources

This study used saved outputs from an integrated antipsychotic pharmacology and lipid-TWAS pipeline (Figure 1). The present Stage 3 analysis reused previously generated long-format receptor-level and drug-level outputs; it did not generate new GWAS, TWAS, receptor-binding, or patient-level clinical data. The antipsychotic drug list was generated from the N05A parent group of the World Health Organization Anatomical Therapeutic Chemical Classification System and Defined Daily Dose Index [22]. The initial list contained 70 drug entries. Nine drugs were subsequently excluded because they lacked a valid defined daily dose, leaving 61 entries for further processing. Fifty-eight drugs were matched to ligand names in the Ki database using fuzzy name matching with a threshold of 80.

Figure 1: Experimental pipeline for prior-sensitivity analysis of the antipsychotic Ki–DDD–TWAS framework(a) Compound and target data harmonization. Antipsychotic entries are parsed from the WHO ATC/DDD Index (group N05A) [21], filtered to those with a valid defined daily dose (DDD), and matched to curated Ki-database ligands by fuzzy name matching (WRatio ≥ 80) [22]. Matched pairs form the drug × receptor table; receptor labels are mapped to HGNC gene symbols and Ki values are aggregated per pair by the minimum across measurements. (b) Per-receptor contribution score. Inverse Ki and the log-transformed DDD are combined into the affinity–dose score, log1p-transformed to give the affinity term. The TWAS p-value is converted to a capped −log10(p) and an exponential scaling factor (β = 0.045); the weighted contribution is the product of affinity, receptor weight and TWAS scale. (c) Dual receptor-weighting schemes. Contributions are computed under a literature-weighted model (fixed metabolic receptor priors, clinically informed reference) and a uniform model (all recognized receptor genes = 1.0, discovery/sensitivity), then summed per drug, normalized to chlorpromazine = 100, and ranked. (d) Transancestral Stage 3 synthesis. Both models are run across five ancestry-stratified GLGC 2021 LDL TWAS datasets [13]; per-receptor contributions are reconstructed from saved outputs and passed through the downstream analytical suite. This figure illustrates the workflow only and does not depict results. Ki: inhibition constant; DDD: defined daily dose; TWAS: transcriptome-wide association study; WHO: World Health Organization; ATC: Anatomical Therapeutic Chemical; HGNC: HUGO Gene Nomenclature Committee; WRatio: weighted ratio; GLGC: Global Lipids Genetics Consortium; LDL: low-density lipoprotein cholesterol; p: p-value; β: beta scaling coefficient; −log10(p): negative base-10 logarithm of the p-value.

Figure 1: Experimental pipeline for prior-sensitivity analysis of the antipsychotic Ki–DDD–TWAS framework(a) Compound and target data harmonization. Antipsychotic entries are parsed from the WHO ATC/DDD Index (group N05A) [21], filtered to those with a valid defined daily dose (DDD), and matched to curated Ki-database ligands by fuzzy name matching (WRatio ≥ 80) [22]. Matched pairs form the drug × receptor table; receptor labels are mapped to HGNC gene symbols and Ki values are aggregated per pair by the minimum across measurements. (b) Per-receptor contribution score. Inverse Ki and the log-transformed DDD are combined into the affinity–dose score, log1p-transformed to give the affinity term. The TWAS p-value is converted to a capped −log10(p) and an exponential scaling factor (β = 0.045); the weighted contribution is the product of affinity, receptor weight and TWAS scale. (c) Dual receptor-weighting schemes. Contributions are computed under a literature-weighted model (fixed metabolic receptor priors, clinically informed reference) and a uniform model (all recognized receptor genes = 1.0, discovery/sensitivity), then summed per drug, normalized to chlorpromazine = 100, and ranked. (d) Transancestral Stage 3 synthesis. Both models are run across five ancestry-stratified GLGC 2021 LDL TWAS datasets [13]; per-receptor contributions are reconstructed from saved outputs and passed through the downstream analytical suite. This figure illustrates the workflow only and does not depict results. Ki: inhibition constant; DDD: defined daily dose; TWAS: transcriptome-wide association study; WHO: World Health Organization; ATC: Anatomical Therapeutic Chemical; HGNC: HUGO Gene Nomenclature Committee; WRatio: weighted ratio; GLGC: Global Lipids Genetics Consortium; LDL: low-density lipoprotein cholesterol; p: p-value; β: beta scaling coefficient; −log10(p): negative base-10 logarithm of the p-value.

The Ki database was an internally curated ligand-receptor dataset identified as KiDatabase_2026-06-22.csv [23]. Ki values were parsed as nanomolar measurements. When multiple measurements were available for the same drug-receptor pair, the minimum observed Ki value was used, corresponding to the strongest recorded binding measurement. Minimum-Ki aggregation was retained because the prespecified pipeline was intended to capture the strongest recorded interaction; a median or geometric-mean Ki sensitivity analysis was not included in the saved Stage 3 run. Defined daily doses were converted to milligrams. The affinity-dose term was then calculated from inverse Ki and the natural logarithm of the milligram dose plus one.

The genetic component consisted of precomputed ancestry-stratified LDL TWAS outputs derived from the GLGC 2021 resource [14]. Five ancestry-defined low-density lipoprotein cholesterol datasets were analyzed: African ancestry (LDL_AFR), East Asian ancestry (LDL_EAS), European ancestry (LDL_EUR), Hispanic ancestry (LDL_HIS), and South Asian ancestry (LDL_SAS). The labels are retained from the supplied data files. The ancestry-specific files contained between 17,620 and 18,057 genes. Each completed pipeline run contained 1,285 drug-receptor pairs. The downstream analysis used only the saved long-format receptor-level files and drug-level risk-score files. No new drug matching, Ki lookup, or TWAS calculation was performed during Stage 3.

Score construction

The upstream affinity-dose score was defined as inverse Ki multiplied by the natural logarithm of defined daily dose in milligrams plus one. Before receptor weighting, this quantity was transformed using the natural logarithm of one plus the affinity-dose score. For each receptor-gene row, the transformed affinity component was multiplied by a receptor weight and by a TWAS scaling factor.

The TWAS scaling factor was calculated as the exponential of 0.045 multiplied by the negative log10 of the TWAS p-value. β = 0.045 was retained from the earlier TWAS-transformation study [19] to provide mild, rather than dominant, amplification. Under this formula, a p-value of 10−5 gives a multiplier of approximately 1.25, while a negative-log10 p-value of 25 gives a maximum multiplier of approximately 3.08. β was not estimated or clinically calibrated in the present LDL analysis. Negative log10 p-values were capped at 25 to prevent extreme amplification. The saved Stage 3 outputs did not include a separate cap-sensitivity sweep, so sensitivity to this cap was not formally estimated here. The saved pipeline also retained a secondary drug-level multiplier based on the maximum absolute TWAS z-score and a nominal-significance multiplier for drug-receptor rows with p-values below 0.05. The nominal p-value threshold was exploratory and was not adjusted for the number of genes, drugs, receptor rows, ancestries, or traits.

Drug-level scores were normalized to chlorpromazine, which was set to 100. This normalization is a reference convention and does not represent a clinical relative risk, incidence rate, probability, or percentage change in LDL. Because the per-receptor transformation used p-values and the secondary multiplier used absolute z-scores, the score emphasized statistical strength and magnitude but did not preserve the direction of the TWAS association. A strong positive and a strong negative LDL association could therefore increase a score in a similar way. The resulting score is therefore direction-agnostic and should not be interpreted as a directional estimate of LDL liability or a causal mediation score.

Two versions of the model were analyzed. The literature-weighted model used predefined metabolic receptor weights. The largest weights were assigned to HRH1, HTR2C, and CHRM3, with values of 1.00, 0.92, and 0.85, respectively. HTR2A received a weight of 0.55. DRD3, DRD2, and DRD4 received weights of 0.18, 0.12, and 0.10. Other receptors were assigned lower weights according to the supplied dictionary.

The uniform model assigned a weight of 1.0 to every recognized receptor gene present in the scored long-format data. It therefore did more than remove differential weighting among the genes already represented in the literature-weight dictionary. It also restored contribution from recognized receptor genes that had no positive literature weight in the weighted model, including genes such as HTR2B and DRD1. The uniform model should consequently be regarded as an equal-weight prior model, not as an absence of pharmacological assumptions.

Stage 3 analytical suite

The Stage 3 analysis reconstructed per-receptor contributions using the saved affinity-dose, TWAS p-value, and gene-mapping fields. Missing p-values were handled using the prespecified fairness-imputation rule, in which missing values among scored genes were replaced with the mean negative log10 p-value available in the relevant run. This rule prevented missingness from automatically setting a receptor contribution to zero, but it also assigned a common value to rows without directly observed p-values. The upstream outputs recorded 210 TWAS-imputed rows in the African, European, Hispanic, and South Asian datasets and 239 in the East Asian dataset. A separate exclusion-based sensitivity analysis was not included in the saved Stage 3 run. This reconstruction was intended to reproduce the contribution formula and did not involve re-matching drugs or re-querying the Ki database.

The analysis first pooled receptor contributions across the five ancestry datasets and calculated the percentage of total contribution attributable to each receptor class. Gene-level contribution percentages were then calculated separately for each ancestry and compared between the weighted and uniform models. The difference was defined as the uniform contribution share minus the weighted contribution share. Directional consistency was defined as the fraction of ancestry datasets showing the same sign of change.

Additional analyses included ancestry-specific rankings of the leading uniform-model genes, Spearman correlations of drug ranks between ancestry datasets, the mean rank and standard deviation of each drug across ancestries, and pairwise Jaccard similarity of the top-10 drug sets. Leave-one-gene-out analyses were performed in the weighted model for HRH1, HTR2C, CHRM3, HTR2A, DRD2, ADRA1A, HTR6, and HTR7. These analyses removed the selected gene’s reconstructed contribution and recalculated the rank from the remaining contribution. They did not fully recompute every secondary drug-level multiplier used in the saved final score.

The high-TWAS/low-weight candidate screen used a strict threshold of an absolute z-score of at least 3.0 and literature weight no greater than 0.15 in the downstream analysis. The Stage 3 synthesis also used a more permissive exploratory threshold of an absolute z-score at least 2.5 with the same weight ceiling. Candidate support was counted by the number of ancestry datasets in which a gene met the threshold. All candidate-gene findings were treated as exploratory model signals rather than as evidence of causal or clinical relevance. A dual-reporting analysis compared each drug’s mean rank under the weighted and uniform models.

Software and reproducibility

All Stage 3 analyses were performed using Python (Python Software Foundation, Wilmington, Delaware, USA) with the open-source packages pandas, NumPy, RapidFuzz, tqdm, and SciPy. The openpyxl package was used to export the spreadsheet. The analyses used saved output files and reconstructed contributions without repeating drug-name matching or receptor lookup. The output directories contained the long-format receptor-level files, drug-level risk files, and supporting Stage 3 summaries.