Work overview

Section 02 of 06

Materials and methods

Artificial intelligence in forensic science: a systematic review. Part I: personal identification

Valentina Bugelli, Francesco Calabrò, Laura Donato, Rossana Cecchi, Jessika Camatti, Marco Di Paolo, and Lorenzo Franceschetti · 2026

Contents

Section 02 of 06

  1. 01Introduction
  2. 02Materials and methods
  3. 03Results
  4. 04Discussion
  5. 05Conclusions
  6. 06Supplementary Information
Text size
Work overview

Section 2 of 6

Materials and methods

Valentina Bugelli, Francesco Calabrò, Laura Donato, Rossana Cecchi, Jessika Camatti, Marco Di Paolo, and Lorenzo Franceschetti · about 3 minutes

Study design and research question

This study was designed as a systematic review aimed at evaluating the application of AI techniques in forensic personal identification. The review was conducted and reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA 2020) guidelines [11].

The study protocol was defined a priori, including the research question, eligibility criteria, and data extraction strategy, in order to ensure methodological transparency and reproducibility. The study protocol was prospectively registered on the Open Science Framework (OSF) platform. A publicly accessible view-only version of the protocol is available at: https://osf.io/39phc/overview?view_only=c223168473564131aec86a269e1beacd.

The aim of this review was to investigate the performance, applicability, and limitations of AI–based approaches used for forensic personal identification, including human identification from imaging data, sex estimation, ancestry or population affinity estimation, and facial recognition in forensic contexts.

The research question was structured according to a modified PICO framework adapted for methodological studies in forensic science: Population/Data: forensic imaging, odontological, anthropological, or other biological datasets; Intervention: AI-based analytical methods; Comparator: conventional forensic methods or human expert assessment when available; Outcome: identification performance metrics (e.g., accuracy, sensitivity, specificity, area under the curve (AUC), equal error rate (EER), and rank-based accuracy).

Eligibility criteria and search strategy

Studies were selected according to the following inclusion criteria: original research articles involving AI-based methods (e.g., machine learning, deep learning, neural networks, or related computational approaches); forensic or medico-legal applications related to personal identification; studies employing imaging, anthropometric, odontological, or other biological data for identification purposes; studies reporting quantitative performance metrics (e.g., accuracy, sensitivity, specificity, AUC, EER, or rank-based accuracy); peer-reviewed articles published in English.

Exclusion criteria were: studies focused exclusively on clinical diagnosis without relevance to forensic identification; editorials, letters, conference abstracts without full data, and opinion papers; studies lacking objective performance metrics. Studies exclusively focused on forensic age estimation were excluded from this review, as this topic is addressed in a separate systematic review within the same three-part series on artificial intelligence applications in forensic science. Conference proceedings and abstracts were excluded due to limited methodological detail and lack of full peer-reviewed publication.

A systematic literature search was conducted in MEDLINE (via PubMed) and Scopus from database inception to the final search date (1 March 2026). The electronic search was supplemented by backward and forward citation tracking in order to identify additional relevant studies. The complete database-specific search strategies are reported in Supplementary Appendix 1. Reference lists of the included studies were also manually screened to identify additional eligible articles.

Study selection and data extraction

All records retrieved from the databases were exported and imported into Zotero reference management software, where duplicate records were identified and removed. Two independent reviewers performed the screening process in two stages: title and abstract screening; full-text assessment for eligibility. Disagreements between reviewers were resolved through discussion and consensus.

Data extraction was conducted using a standardized data extraction form developed prior to the screening phase. The following information was collected from each included study: author(s), year of publication, and journal; dataset characteristics (sample size, population, and type of data); AI model architecture; task type (e.g., classification, identification, or prediction); reported performance metrics; comparator methods (human experts or conventional forensic techniques, when available); and methodological limitations. Data extraction was independently performed by two reviewers and subsequently cross-checked to ensure accuracy and completeness.

Risk of bias, methodological quality assessment, and data synthesis

Methodological quality and risk of bias were evaluated using tools appropriate for diagnostic and predictive studies involving AI. Depending on study design, the following frameworks were applied: QUADAS-2 for diagnostic accuracy studies; PROBAST, adapted for AI prediction models. Particular attention was paid to potential sources of bias specific to AI-based studies, including dataset imbalance, lack of external validation, and possible overfitting of predictive models. In addition, adherence to AI reporting standards (e.g., CLAIM or TRIPOD-AI) was qualitatively assessed when applicable.

Given the expected methodological heterogeneity in datasets, model architectures, and outcome metrics, a narrative synthesis was planned as the primary method of analysis. Included studies were grouped according to their main application domain: dental-based human identification; facial recognition in forensic contexts; sex estimation; ancestry or population affinity estimation. When sufficient methodological homogeneity was identified among studies, quantitative comparison of reported performance metrics was considered.