Section 1 of 6
Introduction
Taiwu Wu, Zhenhua Tang, Shiting Zheng, Hongyan Huang, and Xinzhe Zhu · about 3 minutes
Emerging contaminants (ECs) are increasingly detected in diverse environmental matrices worldwide due to increasing production and the widespread use of synthetic chemicals [1]. Therein, per- and polyfluoroalkyl substances (PFASs), antibiotics (ABs), and endocrine-disrupting chemicals (EDCs) have attracted increasing attention due to their environmental persistence and potent biological effects even at trace concentrations [2,3]. Rivers, as major recipients of pollutants, harbor ECs in both water and sediment, particularly in highly urbanized and industrialized basins [2,4]. Sediments act as long-term sinks and potential sources, releasing ECs back into overlying water under physical, chemical, and biological disturbances [5]. Understanding EC behavior in coupled sediment–water systems under realistic environmental conditions is critical for assessing environmental risks and informing mitigation strategies [6].
Sediment-water partitioning, commonly quantified by the distribution coefficient (log_K_d), is a key determinant of EC mobility and persistence and serves as a foundational parameter in fate-and-transport models and risk assessment [7]. In natural systems, however, ECs rarely reach thermodynamic equilibrium. Measured _K_d values therefore often reflect site-specific pseudo-partitioning coefficients controlled by local environmental conditions rather than universal constants [8]. Direct in situ monitoring of _K_d remains challenging due to ultra-trace concentrations, limited sensing technologies, and the high analytical cost of EC-specific measurements [9,10]. Most available data rely on labor-intensive sampling and laboratory analysis with limited spatiotemporal coverage. These constraints highlight the urgent need for reliable prediction of _K_d from accessible descriptors, particularly across large or data-scarce regions.
Reliable prediction requires an understanding of the factors that govern sediment-water partitioning. Previous field and laboratory studies have identified correlations between _K_d and contaminant molecular descriptors (e.g., octanol–water partition coefficient, log_K_ow), site-specific environmental conditions (e.g., sediment total organic carbon, water salinity, and pH), and basin-scale factors (e.g., proximity to urban areas) [6,[11], [12], [13]]. Conventional water quality indicators have also been reported to be negatively correlated with antibiotic _K_d, suggesting that higher nutrient levels in water may inhibit their partitioning into sediments [14]. However, whether these relationships are transferable across different EC classes remains unclear. Growing evidence suggests that distinct contaminants may exhibit fundamentally different partitioning behaviors depending on their molecular structures and environmental contexts. A systematic framework for understanding how multiple drivers jointly regulate sediment–water partitioning and whether different contaminant classes follow distinct mechanistic regimes is still lacking. Controlled laboratory experiments are informative, but often fail to capture the inherent heterogeneity and dynamic interactions present in natural aquatic systems [[15], [16], [17]].
The increasing availability of field monitoring data offers opportunities to unravel these complex, non-linear relationships using machine learning (ML) [[18], [19], [20]]. ML has been applied to predict the sorption coefficients of pharmaceuticals in soil and sediment and to map the global distribution of heavy metals [[21], [22], [23]]. However, many existing models focus on individual contaminant classes or single-scale mechanisms, with limited interpretability and generalization [24]. Capturing sediment–water partitioning in realistic environments requires integrating heterogeneous drivers operating across multiple spatial scales. Conventional ML architectures, such as artificial neural networks (ANN), typically combine diverse inputs into a single modeling stream, which may obscure hierarchical interactions among molecular, environmental, and basin-scale factors [25]. Hierarchical modeling frameworks that explicitly represent multi-scale drivers may help address this limitation [26]. Complementary molecular dynamics (MD) simulations can further provide microscale insights into interfacial interactions that are difficult to infer from field-derived descriptors alone, thereby strengthening mechanistic interpretation across contaminant classes [27,28].
Here we develop a cross-scale predictive and interpretive framework to investigate the sediment–water partitioning of ECs. The framework uses a multi-branch multi-head attention (MB-MHA) architecture to integrate molecular descriptors at the microscale, water and sediment characteristics at the mesoscale, and basin attributes at the macroscale. We used this framework to: benchmark multi-branch models against conventional ANN architectures; quantify the relative contributions of key micro-, meso-, and macroscale drivers to log_K_d across EC classes; combine model interpretation with MD simulations to examine interfacial mechanisms governing EC partitioning; and demonstrate basin-scale applicability through spatiotemporal forecasting of log_K_d in a rapidly urbanizing river basin. By linking large-scale environmental observations, interpretable modeling, and molecular-level insights, this study characterizes the spatiotemporal heterogeneity and mechanistic classification of sediment–water partitioning across EC classes and provides a framework for predicting EC behavior in complex river basin systems.