Section 4 of 6
Discussion
Valentina Bugelli, Francesco Calabrò, Laura Donato, Rossana Cecchi, Jessika Camatti, Marco Di Paolo, and Lorenzo Franceschetti · about 7 minutes
This systematic review analyzed the current evidence regarding the application of AI techniques in forensic personal identification. A total of 89 studies were included, covering a wide range of forensic tasks such as sex estimation, human identification, ancestry or population affinity estimation, and kinship verification [12–100]. Overall, the findings of this review indicate that AI-based approaches demonstrate high predictive performance across multiple forensic applications. In studies reporting single accuracy metrics, the median accuracy was approximately 91%, suggesting that AI methods may provide valuable support for forensic identification tasks. Sex estimation represented the most frequently investigated application, while deep learning approaches—particularly convolutional neural networks applied to radiological and photographic datasets—were the most commonly used model architectures. Despite these promising results, substantial heterogeneity was observed across datasets, methodological approaches, and evaluation metrics, highlighting the need for careful interpretation of reported performances.
One of the most consistent findings across the included studies was the high predictive performance achieved by AI models in sex estimation tasks. Many studies reported accuracy values exceeding 85–90%, with several deep learning models achieving performance above 95%, consistent with the overall accuracy distribution observed in the included studies (Fig. 2) and with individual studies reporting very high performance in CT-based skeletal analyses [56, 81, 92]. These results are consistent with the well-established sexual dimorphism of specific skeletal regions—particularly the pelvis and skull—which are widely used in forensic anthropology for biological sex estimation [2, 4, 101]. These anatomical differences provide biologically informative features that can be effectively captured by machine learning algorithms. AI-based approaches may therefore enhance traditional morphometric analyses by automatically identifying complex patterns within multidimensional datasets and by reducing potential observer-related variability.
In addition to metric and morphometric approaches, individual skeletal features and pathological variations may also contribute to forensic identification and anthropological reconstruction. Case-based investigations and historical osteological studies have demonstrated how specific skeletal traits, trauma patterns, or anatomical anomalies can assist in reconstructing biological profiles and identifying human remains in both modern forensic contexts and historical investigations [102–104]. These traditional anthropological approaches remain essential in forensic practice and provide valuable reference frameworks for the development and validation of AI-based analytical models.
When comparing model categories, deep learning approaches tended to achieve slightly higher accuracy values than traditional machine learning algorithms. This trend was particularly evident in studies using imaging datasets, where convolutional neural networks demonstrated strong performance in tasks involving radiographs, computed tomography scans, and photographic skeletal images [25, 41, 92]. The ability of deep learning models to automatically extract hierarchical visual features from high-dimensional image data likely contributes to their improved performance in these contexts. However, the difference between deep learning and traditional machine learning models was not always substantial. In several studies relying on structured morphometric measurements, algorithms such as Random Forest, Support Vector Machines, or ensemble learning approaches performed comparably to deep learning models [14, 34, 47]. These findings suggest that the optimal modeling strategy may depend largely on the type and structure of the available data.
Another important observation emerging from this review is the prominent role of medical imaging datasets in AI-based forensic identification research. Computed tomography scans, conventional radiographs, and dental panoramic radiographs represented the most commonly used data sources across the included studies. The increasing availability of digital medical imaging, particularly post-mortem computed tomography (PMCT) in forensic practice, has likely facilitated the rapid adoption of AI-based analytical approaches [5, 7]. Imaging-based datasets offer several advantages, including high anatomical detail, standardized acquisition protocols, and the possibility of extracting large numbers of quantitative features from skeletal structures. As a result, AI-driven image analysis may represent one of the most promising directions for future developments in forensic identification.
Imaging-based approaches have long played a central role in forensic identification, particularly in facial comparison and age progression techniques used in missing persons investigations. These methods rely on the analysis of craniofacial morphology and growth patterns and have been extensively discussed in forensic medicine literature [105–113].
Despite the encouraging performance of AI models reported in the literature, several methodological and practical challenges remain. One of the most notable issues concerns the geographical and demographic distribution of the datasets used for model development. A large proportion of the included studies were conducted on Asian populations, particularly in China, Turkey, Japan, and South Korea. European populations were also represented in several studies, whereas African and South American populations were considerably underrepresented. This imbalance raises important concerns regarding the generalizability of AI models across different populations. Skeletal morphology and biological traits may vary between populations due to genetic, environmental, and developmental factors [19, 49]. Consequently, models trained on population-specific datasets may not perform equally well when applied to individuals from different demographic backgrounds. Addressing this limitation will require the development of more diverse and multi-population datasets in future research.
Another challenge identified in this review relates to methodological heterogeneity across studies. The included studies differed substantially in terms of dataset size, anatomical structures analyzed, model architectures, and evaluation metrics. Sample sizes ranged from fewer than ten individuals in microbiome-based identification studies to more than two hundred thousand radiographic images in large dental datasets. Similarly, performance was reported using a variety of metrics, including accuracy, AUC, sensitivity, specificity, equal error rate, and rank-based identification measures. This heterogeneity limited the possibility of conducting a formal meta-analysis and complicates direct comparisons between studies. Standardization of reporting practices and evaluation metrics would therefore greatly facilitate future evidence synthesis in this rapidly evolving field.
In addition to dataset heterogeneity, several studies exhibited methodological limitations that are commonly encountered in AI-based predictive modeling. These include small training datasets, potential class imbalance, limited transparency regarding model development, and the absence of external validation. External validation using independent datasets is particularly important to assess the robustness and real-world applicability of predictive models. However, many studies relied exclusively on internal validation procedures such as cross-validation or hold-out testing within the same dataset. Without independent validation, reported performance metrics may overestimate the true predictive ability of the models when applied in forensic casework.
The findings of this review are broadly consistent with previous research highlighting the growing role of AI in forensic sciences [5–8]. Recent reviews in forensic medicine and forensic radiology have similarly emphasized the potential of AI-based approaches to assist experts in tasks such as skeletal analysis, dental comparison, and biometric identification. In this context, AI should not be considered a replacement for forensic expertise, but rather a complementary tool capable of supporting expert decision-making and improving analytical efficiency. Integrating AI-based systems with expert interpretation may ultimately enhance both accuracy and reproducibility in forensic identification processes.
The integration of AI into forensic identification also raises important ethical and legal considerations. In medico-legal contexts, algorithmic outputs may influence judicial decisions, making transparency, interpretability, and methodological robustness essential requirements for AI-based systems [5–8]. The integration of complex “black-box” models may limit the ability of forensic experts to explain how specific predictions are generated, potentially creating challenges in legal settings where expert testimony must be clearly justified [114]. In addition, the use of population-specific datasets may introduce potential biases if models are applied to individuals from underrepresented demographic groups [5–8]. Ensuring transparency in model development, rigorous validation procedures, and adherence to emerging reporting guidelines for AI research will therefore be crucial for the responsible implementation of AI in forensic practice.
This review has several strengths. First, it provides a comprehensive synthesis of the available literature on AI-based forensic personal identification across multiple forensic disciplines, including anthropology, odontology, radiology, and biometrics. Second, the review was conducted in accordance with PRISMA guidelines and followed a predefined protocol registered on the Open Science Framework, which increases methodological transparency and reproducibility. Finally, the inclusion of a large number of studies allowed the identification of broad methodological trends and research patterns within the field.
Nevertheless, several limitations should be acknowledged. The review included only peer-reviewed studies published in English, which may introduce a degree of publication bias. In addition, the substantial heterogeneity in datasets, model architectures, and performance metrics prevented quantitative meta-analysis of the results. Finally, the rapidly evolving nature of AI research means that new methodological developments may emerge quickly after the completion of the literature search.
Future research should focus on several key priorities. The development of larger and more diverse datasets, including multi-population skeletal and radiological collections, will be essential to improve the generalizability of AI models. External validation using independent datasets should become a standard practice in the evaluation of forensic AI systems. Furthermore, the integration of explainable AI techniques may help improve transparency and interpretability, which are critical factors in medico-legal contexts where algorithmic decisions may be subject to legal scrutiny. Ultimately, interdisciplinary collaboration between forensic scientists, data scientists, and clinicians will be crucial to ensure the responsible and effective implementation of AI technologies in forensic identification.