Section 2 of 6
Materials and methods
Valentina Bugelli, Francesco Calabrò, Jessika Camatti, Rossana Cecchi, Marco Di Paolo, and Lorenzo Franceschetti · about 3 minutes
Study design and research question
This study was designed as a systematic review aimed at evaluating artificial intelligence (AI) applications in postmortem interval (PMI) estimation. The review was conducted according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA 2020) guidelines, ensuring transparency and reproducibility in study identification, selection, and synthesis [10]. The study protocol was prospectively registered on the Open Science Framework (OSF) to ensure methodological transparency and reproducibility. A publicly accessible view-only version of the protocol is available at: https://osf.io/j3ymz/overview?view_only=f0e11d87cbb14529833330e94ca762e5.
The objective of this review was to investigate the performance, methodological approaches, and forensic applicability of AI-based models developed for estimating the postmortem interval using biological, imaging, biochemical, entomological, or environmental data.
Eligibility criteria
Studies were included if they met the following criteria: original research articles applying AI, machine learning, or deep learning techniques; studies addressing postmortem interval estimation or time-since-death prediction; use of forensic or experimentally simulated postmortem datasets; quantitative evaluation of model performance (e.g., MAE, RMSE, accuracy, correlation coefficients); peer-reviewed publications written in English. Studies involving both human postmortem data and experimentally generated animal models used to simulate decomposition were considered eligible.
Exclusion criteria were: studies using purely statistical or regression-based approaches without AI or machine learning methods; narrative reviews, editorials, conference abstracts without full data, and methodological commentaries; studies focused exclusively on decomposition description without predictive modelling; articles lacking objective performance metrics.
Information sources and search strategy
MEDLINE (via PubMed) and Scopus were searched from database inception to the final search date (1 March 2026), supplemented by backward and forward citation tracking. Database-specific strategies are provided in Supplementary Appendix 1. Manual screening of reference lists from eligible studies was performed to identify additional relevant publications. After duplicate removal, studies were screened independently by two reviewers through: title and abstract screening; full-text evaluation. Disagreements were resolved through consensus discussion.
Data extraction
Data extraction was conducted using a standardized form developed prior to study selection. The following variables were collected: author(s), year, and journal; study design and forensic context; type of data used for PMI estimation (imaging, biochemical, molecular, entomological, environmental, multimodal); sample characteristics (human, animal models, experimental conditions); AI model architecture and training strategy; validation approach (internal or external validation); target outcome (continuous PMI estimation or interval classification); performance metrics (MAE, RMSE, correlation, accuracy); comparator methods (traditional PMI estimation approaches); reported limitations and sources of bias. Data extraction was performed independently by two reviewers and verified for consistency.
Risk of bias and methodological quality assessment
Risk of bias was assessed using tools adapted to AI-based prediction studies, primarily based on the PROBAST framework for predictive modelling studies and QUADAS-2 principles when diagnostic accuracy designs were applicable. Because no standardized risk-of-bias tool currently exists specifically for artificial intelligence–based forensic PMI prediction studies, the assessment criteria were adapted according to methodological characteristics relevant to forensic AI applications.
The evaluation focused on: (i) dataset representativeness; (ii) consideration of environmental variability; (iii) validation strategy (internal versus external validation); (iv) potential risk of overfitting; (v) transparency and reproducibility of AI pipelines. Studies relying exclusively on small experimental datasets, animal-only decomposition models, or internal validation strategies were considered at increased risk of bias regarding generalizability and predictive robustness.
Each study was classified as presenting low, moderate, or high overall risk of bias. A detailed study-level assessment is provided in Supplementary Table S1.