Work overview

Section 03 of 04

Discussion

Fourier Evaluation of Tracings and Acidosis in Labor (FETAL) Framework: An In-Silico Evaluation of a Spectral-Divergence Method

Metadata pending adapter verification · 2026

Contents

Section 03 of 04

  1. 01Introduction
  2. 02Technical report
  3. 03Discussion
  4. 04Conclusions
Text size
Work overview

Section 3 of 4

Discussion

Metadata pending adapter verification · about 8 minutes

The present simulation addresses a deliberately narrow question: whether the FETAL spectral-divergence score behaves coherently and reproducibly under controlled computational conditions. It does not determine whether the score identifies fetal acidemia in clinical recordings, improves prediction beyond established cardiotocographic variables, or adds value beyond clinician interpretation. The findings should therefore be interpreted as evidence of computational identifiability rather than clinical diagnostic validity. Two aspects of the design are central to interpreting the results. First, the reference spectrum was constructed from synthetic reference-like signals generated by the simulation model and was not derived from a cohort of physiologically normal fetuses. Second, the altered spectra were created through controlled redistribution of spectral peaks rather than through a validated biological model of fetal acidemia. Published human studies reporting changes in the distribution of fetal heart-rate variability across frequencies provided the conceptual motivation for examining spectral shape [6-8]. However, the magnitude and precise location of the weak, moderate, and strong shifts were not quantitatively calibrated to effect sizes reported in those studies. The imposed shifts should therefore be understood as mathematical perturbations selected to test the proposed feature under progressively more detectable conditions. Peaks were displaced within prespecified low- and high-frequency regions while approximate total oscillatory power and broad low-/high-frequency allocation were preserved. This design allowed the full normalized spectrum to change without necessarily producing a comparable change in overall standard deviation or the low-/high-frequency power ratio. It consequently tested whether Jensen-Shannon divergence could detect within-band redistribution that scalar summaries might overlook. It does not establish that the simulated patterns reproduce the physiological response to hypoxia or acidemia, and the existence, direction, magnitude, and clinical consistency of such changes must be determined empirically. The method was not guaranteed to discriminate between the simulated groups. Under the null scenario, both test groups arose from the same distribution and the spectral-divergence score remained near chance. The weak-shift scenario was also not reliably distinguishable, indicating that the processing pipeline did not convert every imposed perturbation into apparent discrimination. Performance improved with moderate and strong shifts and declined when the moderate shift was combined with greater artifact burden. These results represent desirable falsification and stress tests: the feature remained uninformative when no usable difference was present, responded progressively when spectral separation increased, and became less reliable when signal quality deteriorated. The strong-shift result must nevertheless be interpreted narrowly. Because the altered signals were generated by relocating spectral peaks and the FETAL spectral-divergence score was designed to quantify full-spectrum divergence, this scenario was favorable to the proposed feature.

The resulting AUROC demonstrates sensitivity to the type of mathematical change the method was designed to detect. It does not demonstrate that acidemic fetuses exhibit this pattern, that the magnitude of the simulated shift is clinically realistic, or that a comparable AUROC would be obtained in humans. The scalar comparators remaining near chance should likewise not be interpreted as evidence that standard deviation or the low-/high-frequency ratio lack clinical value. The simulation was specifically designed to preserve approximate total oscillatory power and broad frequency-band allocation while changing the distribution of power within those bands. The comparators were therefore intentionally less sensitive to the imposed perturbation. Their performance illustrates the distinction between full-spectrum and scalar representations under the stated simulation assumptions, not the comparative clinical superiority of one approach over another.

Jensen-Shannon divergence was used because it provides a symmetric and finite comparison between two normalized spectral distributions on the same frequency grid. These characteristics make the resulting score transparent and straightforward to interpret as a measure of spectral dissimilarity. However, the present study does not establish that Jensen-Shannon divergence is the optimal distributional metric. Future patient-level development should compare it with alternatives such as Kullback-Leibler divergence, Wasserstein distance, cosine similarity, and other measures of spectral shape using prespecified criteria and independent validation. Framed appropriately, this study functions as both an identifiability analysis and a computational unit test. The feature definition, preprocessing pipeline, quality-control rules, spectral estimation procedure, and evaluation scenarios were specified before clinical validation. Null and artifact-stress conditions exposed circumstances in which the method could fail, while the use of independent test signals prevented individual observations from contributing to their own reference template. The documented random-number generation procedure and the availability of the complete simulation code, preprocessing scripts, parameter settings, and replicate-level outputs as supplementary material further support reproducibility. The reported 95% empirical Monte Carlo intervals should not be interpreted as clinical confidence intervals. They represent the 2.5th and 97.5th percentiles of the replicate-specific AUROC values generated under each stated simulation scenario. They therefore summarize variation across repeated realizations of the synthetic experiment rather than sampling uncertainty around diagnostic performance in a human population.

The modest number of replicates also limits the precision with which the tails of these empirical distributions can be characterized. These findings define a hierarchy of evidence for the continued development of the framework. Mathematical coherence asks whether the score is well defined. Computational identifiability asks whether a locked implementation responds to the intended perturbation and remains uninformative under null conditions. Clinical validity asks whether spectral divergence is associated with clinically meaningful outcomes in real intrapartum tracings. Incremental predictive value requires demonstration that the score improves prediction beyond established CTG variables, relevant maternal and fetal characteristics, computerized features, and clinician interpretation. Clinical utility requires further evidence that incorporating the score into decision-making produces a favorable balance of benefits and harms. The present study addresses only the first two levels.

Clinical and public health implications

If subsequently validated in patient-level data, the FETAL spectral-divergence score could potentially serve as an interpretable input to intrapartum clinical decision support. Its intended role would be to complement, rather than replace, visual CTG interpretation. Because it is a transparent, single-valued measure rather than an opaque classifier [14], its contribution to a multivariable model could be examined directly and its relationship with predicted risk could be evaluated for calibration and consistency. However, the present results provide no evidence that adding spectral divergence to established information would improve discrimination, calibration, clinician performance, obstetric intervention rates, or neonatal outcomes. Potential benefits such as reducing unnecessary operative delivery or improving recognition of fetal compromise remain hypothetical. Such claims would require demonstration of reproducible incremental predictive value and clinical utility in appropriately designed human studies. The current simulation establishes only that the feature behaves as intended under specified synthetic conditions, which is a necessary but insufficient step toward clinical use.

Path to clinical validation

Clinical development should use digitized intrapartum CTG recordings linked to paired umbilical-artery blood-gas results [15]. The reference spectrum, preprocessing choices, candidate frequency range, and any model parameters should be developed exclusively within a training cohort. Development and validation must be separated at the patient level so that no windows from the same fetus appear in more than one partition. The spectral-divergence score should first be evaluated as an individual feature and then incorporated into multivariable models containing established CTG characteristics, relevant clinical variables, and, where available, clinician interpretation. A conventional model without spectral divergence should be compared directly with an expanded model containing the feature. This would determine whether the score provides incremental information rather than merely showing an unadjusted association with cord-blood outcomes. Evaluation should include discrimination and calibration [16]. Changes in AUROC alone would be insufficient, particularly if the feature produces only small improvements in ranking. Calibration, clinically relevant sensitivity and specificity, and performance within important subgroups should also be examined. Reporting should follow Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis Plus Artificial Intelligence (TRIPOD+AI) guidance if the feature is incorporated into a clinical prediction model [17]. Umbilical arterial pH is a defensible initial reference outcome because of its association with adverse neonatal outcomes, although it is not synonymous with neurologic injury and does not fully characterize the clinical condition of the newborn [18]. Decision-curve analysis should then be used to quantify net clinical benefit across plausible decision thresholds [19]. External validation across institutions, patient populations, gestational ages, monitoring devices, and clinical practices would be required before implementation. Future studies may also need to examine complementary outcomes such as base deficit, neonatal encephalopathy, resuscitation requirements, and short-term neonatal morbidity while avoiding data-driven redefinition of endpoints after model evaluation.

Limitations

The generative model was intentionally simplified and does not reproduce the full complexity of real intrapartum cardiotocography. Clinical recordings are affected by nonstationarity, temporal relationships with uterine contractions, fetal sleep and behavioral states [20], gestational-age variation, stage and duration of labor, fetal movement, maternal medications, oxytocin exposure, neuraxial anesthesia, progressive physiological deterioration, and changing baseline and deceleration patterns. Real recordings also contain more complex artifact structures than those represented here. These include prolonged signal loss, abrupt transitions between fetal and maternal heart rates, repeated interpolation, transducer displacement, variable sampling and smoothing algorithms, and differences among external Doppler devices, fetal scalp electrodes, software platforms, and institutions. Such factors could alter normalized spectra, increase within-class variability, or generate divergence unrelated to fetal physiology. The synthetic reference template was not a physiological reference, and the imposed shifts were not empirically estimated signatures of acidemia. The simulation therefore cannot establish which spectral patterns, if any, characterize compromised fetuses. The artifact-stress scenario represented only one prespecified combination of increased noise, missingness, decelerations, and maternal contamination and cannot encompass the full range of signal-quality problems encountered clinically. The number of Monte Carlo replicates was modest, and the empirical intervals describe only variability under the chosen generative assumptions. In addition, although the simulation compared spectral divergence with two scalar summaries, it did not compare the feature with comprehensive computerized CTG models, expert interpretation, or alternative spectral-distance metrics. Most importantly, no human tracings or clinical outcomes were analyzed. The reported AUROCs are therefore not estimates of diagnostic accuracy in patients. No conclusions regarding prediction of fetal acidemia, incremental value beyond established CTG assessment, clinical effectiveness, or improvement in intrapartum outcomes should be drawn from the present study.