Work overview

Section 03 of 05

Results and discussion

Experimentally derived biomimetic chromatographic descriptors for drug-induced phospholipidosis liability prediction

Chrysanthos Stergiopoulos and Valko Klara · 2026

Contents

Section 03 of 05

  1. 01Introduction
  2. 02Experimental
  3. 03Results and discussion
  4. 04Conclusions
  5. 05Supplementary material
Text size
Work overview

Section 3 of 5

Results and discussion

Chrysanthos Stergiopoulos and Valko Klara · about 15 minutes

Dataset partitioning and distribution

The training and external test sets (52 and 13 compounds, respectively) showed comparable descriptive statistics, indicating a balanced partitioning of the dataset (Table S4). The training and test sets exhibited similar central tendency and dispersion (mean pEC₅₀: 4.63 vs. 4.51; median: 4.85 vs. 4.73; SD: 0.706 vs. 0.765), as well as overlapping activity ranges (3.10 to 5.87 for training and 3.00-5.41 for test). The distribution of pEC₅₀ values (Figure S1) further supports this, showing consistent coverage across the activity spectrum without apparent major gaps or clustering. Importantly, the class distribution of phospholipidosis (PLD) induction was maintained between the two subsets (Table S4), with similar proportions of non-inducers (25.0 vs. 30.8 %), weak/moderate inducers (38.5 vs. 38.5 %), and strong inducers (36.5 vs. 30.8 %). Likewise, the distribution of ionization states remained consistent across the training and test sets (Figure S2), ensuring that key physicochemical characteristics relevant to lysosomal accumulation were not biased by the split.

The representativeness of the split was further evaluated through principal component analysis (Figure 1). The first two principal components accounted for a substantial proportion of the total variance (PC1: 53.3 %, PC2: 20.7 %), indicating that the projection captures the major sources of variability in the dataset. The PCA plot shows that the test set compounds are well distributed within the chemical space defined by the training set, with no systematic separation or extrapolation beyond the training set's boundaries.

Figure 1.: PCA plot illustrating the chemical-space coverage of the training and external test sets. Blue circles represent training compounds and orange squares represent external test compounds. The broader dashed confidence ellipse corresponds to the training set, whereas the smaller dashed confidence ellipse corresponds to the external test set

Figure 1.: PCA plot illustrating the chemical-space coverage of the training and external test sets. Blue circles represent training compounds and orange squares represent external test compounds. The broader dashed confidence ellipse corresponds to the training set, whereas the smaller dashed confidence ellipse corresponds to the external test set

The strong overlap between the two subsets, as illustrated by the clustering and confidence ellipses, supports that the external test compounds are located within the descriptor space sampled by the training set, reducing the likelihood of extrapolative predictions.

Multiple linear regression models for phospholipidosis prediction

A series of MLR models was developed to evaluate the ability of conventional physicochemical and biomimetic chromatographic descriptors to predict PLD potency, expressed as pEC₅₀. The performance of representative models is summarized in Table 1, while the full set of developed models and statistical parameters is provided in Supplementary Table S5.

Model | Equation | R 2 | Q 2 cv | Q 2 ext | RMSEP
CHI IAM | pEC₅₀ = 3.077 + 0.038·CHI IAM | 0.699 | 0.657 | 0.821 | 0.315
CHI IAM + fneg | pEC₅₀ = 3.22 + 0.035·CHI IAM − 0.346·fneg | 0.724 | 0.689 | 0.823 | 0.313
CHI IAM + log kAGP | pEC₅₀ = 3.58 + 0.022·CHI IAM + 0.317·log kAGP | 0.740 | 0.700 | 0.766 | 0.360
log D7.4 | pEC₅₀ = 4.33 + 0.182·log D7.4 | 0.365 | 0.301 | 0.676 | 0.423
log D7.4 + fpos | pEC₅₀ = 3.68 + 0.177·log D7.4 + 0.878·fpos | 0.579 | 0.479 | 0.761 | 0.364
log kAGP | pEC₅₀ = 4.36 + 0.655·log kAGP | 0.677 | 0.648 | 0.539 | 0.506
log kAGP + fpos | pEC₅₀ = 4.06 + 0.594·log kAGP + 0.444·fpos | 0.726 | 0.665 | 0.586 | 0.479

According to these predefined validation criteria, the CHI IAM-based models satisfied the main requirements for predictive interpretation. The CHI IAM model showed strong internal and external performance, with _Q_2cv = 0.657 and _Q_2ext = 0.821, while the CHI IAM + _f_neg model showed slightly improved performance, with _Q_2cv = 0.689 and _Q_2ext = 0.823. Both models also showed acceptable RMSEP values relative to the dataset's pEC₅₀ range. The CHI IAM + log _k_AGP model showed the highest goodness of fit, but its external predictive performance was lower than that of CHI IAM alone and CHI IAM + _f_neg. By contrast, some conventional descriptor models, such as log _D_7.4 alone, did not meet the same internal validation criterion and were therefore interpreted primarily as baseline comparators rather than preferred predictive models.

Initial models based on conventional lipophilicity descriptors, including log P and log D, showed moderate predictive performance, indicating that bulk lipophilicity alone is insufficient to capture PLD potency in the studied dataset. The inclusion of the positively charged fraction (_f_pos) improved these conventional models, which is consistent with the established role of cationic amphiphilic drug-like behavior in PLD induction [4-8]. Cationic amphiphilic drugs typically combine a lipophilic domain with protonatable basic functionality, enabling membrane permeation, subsequent protonation, and lysosomal sequestration in acidic intracellular compartments [5,6]. Thus, the improvement observed after inclusion of _f_pos is consistent with the contribution of ionization state to lysosomal accumulation and PLD liability.

By comparison, biomimetic chromatographic descriptors showed stronger predictive performance than conventional lipophilicity descriptors in this dataset. The CHI IAM descriptor, which reflects interaction with phospholipid-like surfaces, showed a strong association with PLD potency (_R_2 = 0.699, _Q_2ext = 0.821) and outperformed the single conventional descriptors evaluated in Table 1. This observation is mechanistically plausible because PLD involves accumulation and perturbation within phospholipid-rich intracellular compartments, whereas octanol/water partitioning provides only a simplified representation of molecular lipophilicity [4-6,20]. IAM stationary phases provide an experimentally accessible phospholipid-like environment, and IAM retention has been used as a surrogate for membrane affinity, membrane partitioning, and related ADMET properties [21,23,24]. Previous IAM-based PLD studies further support the relevance of chromatographic membrane-affinity measurements for assessing phospholipidosis [14-16]. In this context, the strong performance of CHI IAM suggests that experimentally measured membrane affinity captures PLD-relevant information not fully represented by calculated log P or log D.

The addition of charge-related information further refined the CHI IAM-based models. The CHI IAM + _f_neg model showed the highest external predictive performance among the representative MLR models (_Q_2ext = 0.823; RMSEP = 0.313), although the improvement over CHI IAM alone was small. This suggests that ionization-related descriptors may provide complementary information, but that CHI IAM already captures a substantial part of the PLD-relevant signal. Although PLD is classically associated with cationic amphiphilic compounds, the broader ionization profile of a molecule may still influence membrane association, intracellular distribution, and apparent PLD potency.

The CHI IAM + log _k_AGP model provided the highest goodness-of-fit among the representative models (_R_2 = 0.740), but its external predictivity was lower than that of CHI IAM alone and CHI IAM + _f_neg (_Q_2ext = 0.766 vs. 0.821-0.823). This pattern indicates that log _k_AGP may capture additional distribution-related information within the training set, but its added predictive value was not consistently reflected in the external test set. From a mechanistic perspective, AGP binding should not be interpreted as a direct driver of PLD. Rather, log _k_AGP can be viewed as an experimentally derived descriptor that reflects physicochemical and distributional features common to many basic and amphiphilic drugs [22,36]. Therefore, its contribution is best interpreted as a complementary distribution-related signal rather than as an independent mechanistic cause of phospholipid accumulation.

Because IAM retention and AGP binding are both influenced by amphiphilicity, ionization, and molecular size, these descriptors are not expected to be fully independent. This behaviour is consistent with the overlapping physicochemical determinants of membrane affinity and plasma protein binding in ADMET-related descriptor sets [21-24,36]. VIF analysis (Supplementary material Table S5) indicated acceptable VIF values, suggesting that multicollinearity was not problematic in the retained models. To further assess the effect of descriptor interdependence, partial least squares regression was applied as a complementary analysis (Supplementary material Table S6). The PLS results indicated that a single latent component captured most of the predictive variance, whereas additional components did not substantially improve model performance. These findings support the interpretation that the predictive information is partly shared across related distributional descriptors, and that CHI IAM provides a practical experimental surrogate for this membrane-affinity-dominated descriptor space.

The comparatively strong performance of log _k_AGP alone (_R_2 = 0.677) and the weaker performance of log kHSA alone (_R_2 = 0.313) are consistent with the tendency of AGP to bind many basic drugs, whereas HSA more commonly contributes to the binding of neutral and acidic compounds [22,36]. Although incorporating _f_pos and _f_neg improved the performance of log _k_HSA-based models, these models remained less predictive than the strongest CHI IAM-based models. This supports the conclusion that membrane-affinity information is central to the present modelling framework, whereas protein-binding descriptors provide more limited, model-dependent complementary information.

This interpretation is illustrated in Figure 2, which shows the agreement between observed and predicted pEC₅₀ values for the CHI IAM + log _k_AGP model in the training and external test sets.

Figure 2.: Observed versus predicted pEC₅₀ values for the MLR model based on CHI IAM and log kAGP. Training and external test sets are shown. The solid line represents the ideal y = x relationship. Model performance: R2 = 0.740, Q2cv = 0.700, Q2ext = 0.766 and RMSEP = 0.360

Figure 2.: Observed versus predicted pEC₅₀ values for the MLR model based on CHI IAM and log kAGP. Training and external test sets are shown. The solid line represents the ideal y = x relationship. Model performance: R2 = 0.740, Q2cv = 0.700, Q2ext = 0.766 and RMSEP = 0.360

The diagnostic plots further support the stability of the selected biomimetic model. In Figure 3, the residuals of the CHI IAM + log _k_AGP model are distributed around zero without obvious systematic curvature, suggesting no major visual evidence of bias.

Figure 3.: Residuals plot for the CHI IAM + log kAGP MLR model. Residuals, calculated as observed − predicted pEC₅₀, are plotted against predicted pEC₅₀ values for both training and test compounds

Figure 3.: Residuals plot for the CHI IAM + log kAGP MLR model. Residuals, calculated as observed − predicted pEC₅₀, are plotted against predicted pEC₅₀ values for both training and test compounds

Figure 4 shows that most compounds fall within the model applicability domain, with leverage values below the warning threshold and standardized residuals within ±3. Together with the VIF analysis and statistically significant regression coefficients reported in Supplementary material Table S5, these diagnostics support the use of the model for interpretation within the chemical space covered by the dataset.

Figure 4.: Williams plot for the CHI IAM + log kAGP MLR model. Standardized residuals are plotted against leverage values for training and external test compounds. The horizontal dashed lines indicate the ±3 standardized residual limits, while the vertical dashed line represents the leverage warning threshold. Cimetidine and Ribavirin exceeded the leverage threshold and are labelled in the plot

Figure 4.: Williams plot for the CHI IAM + log kAGP MLR model. Standardized residuals are plotted against leverage values for training and external test compounds. The horizontal dashed lines indicate the ±3 standardized residual limits, while the vertical dashed line represents the leverage warning threshold. Cimetidine and Ribavirin exceeded the leverage threshold and are labelled in the plot

The Williams plot identified two compounds outside the leverage warning threshold: Cimetidine and Ribavirin. Importantly, neither compound exceeded the ±3 standardized residual limits, indicating that they should be regarded as high-leverage compounds rather than response outliers. Ribavirin is a highly polar nucleoside analog with very low biomimetic membrane and AGP retention, placing it at the low-affinity edge of the descriptor space and outside the main region occupied by typical cationic amphiphilic PLD inducers. Cimetidine is also a comparatively polar and structurally atypical compound, combining imidazole, cyanoguanidine, and thioether functionalities, with relatively low membrane affinity but measurable PLD activity. These structural features may explain why both compounds occupy influential positions in the model space. However, their residuals remained within acceptable limits, suggesting that they define the edge of the model's applicability domain rather than undermining it.

The comparison between descriptor classes indicates that PLD potency is not adequately described by conventional lipophilicity metrics alone. The stronger performance of CHI IAM-based models is consistent with the involvement of membrane-affinity and distribution-related processes in PLD induction. Positive charge and AGP-related retention may modulate this behaviour, but their contributions appear secondary or context-dependent compared with the membrane-affinity signal captured by CHI IAM. This interpretation is consistent with the accepted mechanistic framework of PLD, in which lysosomal sequestration, drug-phospholipid interactions, and altered phospholipid turnover arise from the interplay of amphiphilicity, ionization, and membrane partitioning [4-7].

Mechanistically, these findings are consistent with lysosomal trapping of cationic amphiphilic compounds. After passive diffusion across cellular membranes, protonatable molecules can become increasingly ionized in the acidic lysosomal environment, reducing their ability to diffuse back across the membrane and promoting intracellular accumulation [5,6]. The resulting enrichment of amphiphilic basic compounds in lysosomes can favour interactions with phospholipid-rich structures, thereby contributing to the disruption of lysosomal lipid homeostasis [4-6]. Within this framework, CHI IAM should be interpreted as an experimentally derived surrogate of membrane affinity rather than as a direct measurement of lysosomal accumulation. Its predictive performance supports its use as a practical descriptor for early PLD liability assessment, particularly when combined with appropriate validation and applicability-domain analysis.

Ordinal regression analysis of phospholipidosis classification

Following the development of continuous models for PLD potency using MLR, an ordinal regression approach was applied to evaluate whether the same descriptors could classify compounds into ordered PLD potency categories. This analysis complements the continuous regression framework by focusing on categorical risk stratification, which is particularly relevant for early-stage screening. In this context, reducing false-negative classifications is important, as a practical objective in early discovery is to flag compounds with elevated PLD liability for further evaluation or prioritization. The ordinal regression analysis was not intended to replace the continuous MLR equations or to provide a more readily transferable predictive equation; rather, it was used as a complementary analysis to examine whether the same descriptors could support categorical PLD risk stratification.

As summarized in Table 2, the CHI IAM model achieved the most balanced ordinal classification performance among the representative models, with a test accuracy of 0.769, a macro-F1 score of 0.778, and a macro-AUC of 0.923. This result is consistent with the full ordinal model comparison reported in Supplementary Table S7 and with the continuous MLR analysis, indicating that the same descriptor hierarchy was preserved when PLD liability was expressed as ordered potency classes rather than as continuous pEC₅₀ values. CHI IAM remained statistically significant in the relevant ordinal models, whereas additional descripttors such as _f_neg or log _k_AGP did not consistently improve classification performance or provide significant additional contributions. Protein-binding and conventional lipophilicity descriptors showed useful but less consistent classification behaviour, particularly when evaluated by macro-F1 and class assignment accuracy. Therefore, the ordinal models are interpreted as supportive risk-stratification tools rather than as an independent source of mechanistic evidence.

Model class | Model | Test accuracy | Test macro F1 | Test AUC (macro OVR) | McFadden’s R2
Membrane-binding | CHI IAM | 0.769 | 0.778 | 0.923 | 0.440
Protein-binding | log kAGP | 0.769 | 0.761 | 0.861 | 0.377
Conventional | log D7.4 | 0.692 | 0.689 | 0.915 | 0.169

Using the predefined ordinal-model criteria, the CHI IAM model showed the most balanced validation profile, with McFadden’s pseudo-_R_2 = 0.440, macro-F1 = 0.778, and macro-AUC = 0.923. These values indicate meaningful improvement over the null model, balanced class-level performance, and strong one-vs-rest discrimination. The log _k_AGP and log _D_7.4 models showed useful but less consistent profiles, particularly when macro-F1 and McFadden’s pseudo-_R_2 were considered together.

The confusion matrix for the CHI IAM model (Figure 5) provides additional insight into classification behaviour. Misclassifications occurred mainly between adjacent PLD classes rather than between distant categories, consistent with the ordered nature of the endpoint. The corresponding training-set confusion matrix is provided in Figure S3 and shows a similar pattern, with no direct confusion between non-inducers and strong inducers.

Figure 5.: Confusion matrix for the ordinal regression model based on CHI IAM, evaluated on the external test set

Figure 5.: Confusion matrix for the ordinal regression model based on CHI IAM, evaluated on the external test set

The one-vs-rest ROC analysis (Figure 6) showed stronger discrimination for the non-inducer and strong-inducer classes than for the weak/moderate inducer class. The corresponding training-set ROC curves are provided in Figure S4. These results support the use of ordinal regression as a complementary classification analysis, while the main mechanistic interpretation remains based on the continuous MLR models.

Figure 6.: Multiclass receiver operating characteristic curves for the CHI IAM ordinal regression model evaluated on the external test set. ROC curves were generated using a one-vs-rest approach, in which each PLD class was considered separately as the positive class and the remaining two classes were combined as the negative group. The class-specific AUC values therefore describe the model's ability to discriminate non-inducers, weak/moderate inducers, and strong inducers from all other compounds

Figure 6.: Multiclass receiver operating characteristic curves for the CHI IAM ordinal regression model evaluated on the external test set. ROC curves were generated using a one-vs-rest approach, in which each PLD class was considered separately as the positive class and the remaining two classes were combined as the negative group. The class-specific AUC values therefore describe the model's ability to discriminate non-inducers, weak/moderate inducers, and strong inducers from all other compounds

A broader comparison across descriptor families (Figure 7) further illustrates that the ordinal classification results follow the same trend observed in the continuous MLR models, with CHI IAM showing the most balanced performance among the representative descriptor classes.

Figure 7.: Comparative performance of ordinal regression models based on membrane-binding, protein-binding, and conventional descriptors (significant coefficients only). Test accuracy and macro F1 score are shown for each model

Figure 7.: Comparative performance of ordinal regression models based on membrane-binding, protein-binding, and conventional descriptors (significant coefficients only). Test accuracy and macro F1 score are shown for each model

The ordinal regression analysis confirmed the same descriptor hierarchy observed in the continuous MLR models. CHI IAM provided the most consistent performance, whereas protein-binding and conventional lipophilicity descriptors showed more model-dependent behaviour. The model coefficient estimates and class thresholds reported in Supplementary Tables S8 and S9 further support the interpretability of the ordinal modelling framework within the chemical space represented by the dataset. Therefore, the ordinal analysis was retained as a complementary risk-classification exercise, while the main mechanistic interpretation was consolidated with the MLR-based discussion above. These models should be viewed as early-stage prioritization tools rather than replacements for cell-based PLD assays.