Section 3 of 6
Results
Valentina Bugelli, Francesco Calabrò, Laura Donato, Rossana Cecchi, Jessika Camatti, Marco Di Paolo, and Lorenzo Franceschetti · about 7 minutes
Study selection
The literature search identified a total of 1,266 records through database searching, including 396 records from PubMed/MEDLINE and 870 records from Scopus. After removal of 185 duplicate records and one retracted article, 1,080 records remained for title and abstract screening. During the screening phase, 863 records were excluded based on title and abstract because they did not meet the eligibility criteria. The full texts of 217 articles were then sought for retrieval. Of these, 14 reports could not be retrieved, leaving 203 articles assessed for full-text eligibility.
Following full-text evaluation, 114 articles were excluded for the following reasons: not related to forensic personal identification (n = 107); not focused on AI-based personal identification (n = 5); clinical setting without forensic relevance (n = 1); ethical commentary article (n = 1). Ultimately, 89 studies met the inclusion criteria and were included in the qualitative synthesis of this systematic review [12–100]. The study selection process is summarized in Fig. 1 (PRISMA flow diagram).

Fig. 1: PRISMA Flow diagram of the systematic review
Characteristics of included studies
The included studies were published between 2012 and 2026 and covered multiple forensic applications of AI, including sex estimation, ancestry estimation, human identification, and kinship verification.
The studies were conducted across several geographic regions. The majority of studies originated from Asia, particularly China, Japan, Turkey, and South Korea, followed by Europe, South America, and North America. Sample sizes varied substantially across studies, ranging from 10 individuals in microbiome-based identification research to over 200,000 radiographic images in large dental datasets.
The main characteristics of the studies included in this systematic review are summarized in Table 1, while a detailed description of each study is provided in Supplementary Table S1.
Characteristic | Category | Studies n (%)
Forensic task | Sex estimation | 56 (63.0)
| Human identification | 14 (15.7)
| Ancestry estimation | 9 (10.1)
| Multi-task prediction | 6 (6.7)
| Kinship verification | 4 (4.5)
Data type | Computed tomography-based datasets | 34 (38.2)
| Conventional radiographs | 21 (23.6)
| Dental panoramic radiographs | 10 (11.2)
| Photographic skeletal images | 12 (13.5)
| Other datasets | 12 (13.5)
AI model category | Deep learning | 47 (52.8)
| Traditional machine learning | 28 (31.5)
| Hybrid approaches | 14 (15.7)
Data types and imaging modalities
A wide range of anatomical structures and imaging modalities were analyzed across the included studies. The most frequently used data sources included: computed tomography (CT) scans, particularly cranial and pelvic CT datasets; dental panoramic radiographs (orthopantomograms); conventional radiographs (X-ray images); three-dimensional skeletal reconstructions; photographic images of skeletal elements; facial and ear biometric images; genetic and microbiome profiles.
Among imaging-based approaches, CT-derived measurements and radiographic datasets represented the most common data types, reflecting their widespread use in forensic anthropology and forensic identification research.
AI approaches
A wide variety of AI models were employed across the included studies. The most commonly used algorithms included: Convolutional Neural Networks (CNNs); Artificial Neural Networks (ANNs); Random Forest (RF); Support Vector Machines (SVM); k-Nearest Neighbors (KNN); Logistic Regression (LR); Gradient Boosting and XGBoost models. Deep learning architectures such as ResNet, EfficientNet, GoogLeNet, and VGG-based networks were frequently used in studies involving medical imaging. Several studies also implemented ensemble learning approaches, combining multiple machine learning algorithms to improve predictive performance.
Imaging-based datasets were predominantly analyzed using deep learning architectures, particularly convolutional neural networks, whereas studies relying on morphometric or osteometric measurements more frequently employed traditional machine learning algorithms such as Random Forest or Support Vector Machines.
Forensic tasks and model performance
The majority of studies focused on sex estimation, which represented the most common forensic application across the literature. Other investigated tasks included human identification, ancestry or population affinity estimation, kinship verification, and multi-task prediction frameworks combining sex and age estimation.
Overall, AI approaches demonstrated high predictive performance across most forensic applications. In sex estimation studies, reported accuracies frequently exceeded 85–90%, with several deep learning models achieving accuracy values above 95%. The highest reported performances were observed in CT-based deep learning studies, particularly in pelvic and cranial analyses, where accuracies approached 100% in some datasets. Similarly, identification systems based on dental radiographs or biometric features showed high rank-based identification accuracy, often exceeding 90%. However, performance varied depending on the anatomical structure analyzed, dataset size, and population characteristics, highlighting the importance of population-specific validation.
Across studies reporting single-value accuracy metrics, the median accuracy was 91.4%, with an interquartile range from 88.9% to 95.0%. The distribution of accuracy values across studies is illustrated in Fig. 2. When stratified by model category, deep learning approaches showed slightly higher median accuracy values than traditional machine learning algorithms, although substantial overlap between distributions was observed (Fig. 3).

Fig. 2: Distribution of accuracy values reported in artificial intelligence models for sex estimation. Boxplot representing the distribution of reported accuracy values across studies providing explicit numeric accuracy outcomes. The central line represents the median, the box indicates the interquartile range (IQR), and whiskers represent the minimum and maximum values. Individual points correspond to accuracy values reported by each included study

Fig. 3: Comparison of accuracy distributions between traditional machine learning and deep learning models used for sex estimation. Boxplots showing the distribution of reported accuracy values across studies using traditional machine learning algorithms (n = 5) and deep learning models (n = 14). The central line represents the median, boxes indicate the interquartile range (IQR), and whiskers represent the minimum and maximum values. Individual points correspond to accuracy values reported by each included study
The heterogeneity of datasets, evaluation metrics, and methodological approaches limited the possibility of conducting a formal meta-analysis.
Geographical distribution of studies and population representation
The included studies demonstrated a broad geographical distribution, although a clear concentration of research activity was observed in specific regions. Most studies were conducted in Asia, particularly in China, Turkey, Japan, South Korea, and Thailand, which together accounted for the largest proportion of published datasets. A substantial number of studies also originated from Europe, including populations from France, Portugal, Italy, Greece, Spain, and Bulgaria. Additional contributions were reported from North America, South America, and Africa, although these regions were comparatively less represented.
Several studies used population-specific skeletal collections or hospital-based imaging datasets, often reflecting the demographic composition of the local population. Despite this geographic diversity, the overall distribution of datasets revealed uneven population representation, with a predominance of studies conducted on East Asian and Turkish populations. In contrast, African, South American, and multi-ancestry populations remain underrepresented.
Temporal trends in AI applications
An increasing trend in the application of AI to forensic anthropology and identification was observed over time. The earliest studies included in this review date back to 2012, when machine learning approaches were primarily applied to biometric identification tasks, such as ear recognition or basic morphometric classification. Between 2015 and 2019, the number of publications gradually increased, with the introduction of more advanced machine learning techniques applied to skeletal morphometrics and radiological datasets. During this period, algorithms such as Support Vector Machines, Random Forest, and Artificial Neural Networks were commonly used. A marked increase in publications was observed after 2020, coinciding with the rapid adoption of deep learning architectures and the growing availability of large medical imaging datasets. Recent studies published between 2023 and 2026 frequently incorporated advanced architectures, including ResNet, EfficientNet, Transformer-based networks, and hybrid deep learning frameworks, often achieving higher predictive performance than traditional machine learning approaches.
Overall, these trends indicate a clear shift from traditional morphometric statistical models toward data-driven AI methods, reflecting broader developments in medical imaging and computational anthropology.