Work overview

Section 03 of 06

Results

Generative artificial intelligence in forensic medicine: a pilot study on AI-simulated medico-legal reports in healthcare liability cases

Federica Ministeri, Massimiliano Esposito, Martina Francaviglia, Lucio Di Mauro, Grazia Giulia Pantè, Monica Salerno, Cristoforo Pomara, and Francesco Sessa · 2026

Contents

Section 03 of 06

  1. 01Introduction
  2. 02Materials and methods
  3. 03Results
  4. 04Discussion
  5. 05Conclusions
  6. 06Electronic Supplementary Material
Text size
Work overview

Section 3 of 6

Results

Federica Ministeri, Massimiliano Esposito, Martina Francaviglia, Lucio Di Mauro, Grazia Giulia Pantè, Monica Salerno, Cristoforo Pomara, and Francesco Sessa · about 6 minutes

The evaluation of AI-generated medico-legal reports was conducted according to the predefined multi-level qualitative framework described in the Methods section. Results are presented with reference to three main domains: (1) coherence and quality of medico-legal reasoning, (2) quality of clinical synthesis, (3) characteristics and quality of bibliographic references, and (4) concordance of permanent biological impairment estimates in liability-confirmed cases.

Descriptive statistics were generated and stratified by case category (with vs. without confirmed healthcare liability), as well as aggregated across the entire dataset. This framework enabled an assessment of the overall reliability of AI-generated reports, the functional consistency of the two models, and the scientific appropriateness of the citations automatically inserted by the AI.

Across the 10 selected cases, AI-generated medico-legal reports were compared with corresponding human-authored evaluations. Argumentative coherence and concordance were assessed using a 1–10 qualitative score. Detailed case-by-case scoring is provided in _Supplementary Material _5.

In the five cases classified as non-liability, the hospital-oriented GPT model achieved higher coherence scores, with a mean value of 8.0, compared with a mean score of 5.4 for the patient-oriented model.

Conversely, in the five cases involving confirmed healthcare liability, the patient-oriented GPT demonstrated higher coherence, with a mean score of 5.2, whereas the hospital-oriented model showed a lower mean score of 3.6.

Considering the entire case series, the overall mean coherence score was 5.7 for the hospital-oriented model and 5.2 for the patient-oriented model. These results indicate a consistent alignment between each model’s predefined evaluative orientation and the medico-legal conclusions reached in human-authored reports (Fig. 1).

Fig. 1: Qualitative statistical analysis of the argumentative coherence of AI-generated medico-legal reports (score range: 1–10)

Fig. 1: Qualitative statistical analysis of the argumentative coherence of AI-generated medico-legal reports (score range: 1–10)

Quality of clinical synthesis

The quality of clinical synthesis generated by both GPT models was evaluated using a 1–3 scale, focusing on clarity, completeness, and accuracy in reconstructing the clinical course. Detailed case-by-case scoring is provided in _Supplementary Material _6.

The patient-oriented model achieved a mean synthesis score of 2.6, while the hospital-oriented model achieved a mean score of 2.5. Scores were generally stable across cases and did not differ substantially between liability and non-liability groups. Overall, both models demonstrated a comparable ability to summarize and organize clinical documentation in a coherent manner (Fig. 2).

Fig. 2: Qualitative statistical analysis of AI-generated clinical synthesis (score range: 1–3)

Fig. 2: Qualitative statistical analysis of AI-generated clinical synthesis (score range: 1–3)

Analysis of bibliographic references

The bibliographic references generated by the two GPT models were analyzed both quantitatively and qualitatively and compared with those reported in human-authored medico-legal reports. Detailed case-by-case scoring is provided in _Supplementary Material _7 and 8.

The mean number of references generated per case was 5.6 for the patient-oriented model and 6.7 for the hospital-oriented model. In comparison, human-authored reports included a mean of 0.2 references per case in patient-side reports and 7 references per case in hospital-side reports (Fig. 3).

Fig. 3: Quantitative statistical analysis of bibliographic sources

Fig. 3: Quantitative statistical analysis of bibliographic sources

Qualitative analysis of AI-generated references revealed differences in formal accuracy between the two models. The hospital-oriented GPT produced references with correct authorship in 88.05% of cases, correct journal attribution in 86.56%, and correct publication year in 85.07%. The corresponding values for the patient-oriented model were lower, with authorship accuracy of 64.2%, journal correctness of 69.64%, and publication year accuracy of 69.64%. Valid DOI or bibliographic link availability was observed in 56.7% of references generated by the hospital-oriented model and in 48.21% of those generated by the patient-oriented model (Fig. 4).

Fig. 4: Qualitative statistical analysis of the formal accuracy of authorship, journal, publication year, and DOI/link of AI-generated references

Fig. 4: Qualitative statistical analysis of the formal accuracy of authorship, journal, publication year, and DOI/link of AI-generated references

Regarding relevance to the specific clinical cases, references generated by the patient-oriented model were classified as highly relevant (score 3) in 48.2% of cases, relevant (score 2) in 7.14%, partially relevant (score 1) in 5.35%, and non-relevant (score 0) in 39.2%. In contrast, the hospital-oriented model produced a higher proportion of relevant references, with 49.25% classified as highly relevant, 20.89% as relevant, 13.4% as partially relevant, and 16.41% as non-relevant. The mean relevance score (maximum 3) was 1.64 for the patient-oriented model and 2.02 for the hospital-oriented model.

Across both models, a subset of generated references was found to be non-existent or unverifiable, resulting in a score of zero across all evaluated parameters (Figs. 5 and 6).

Fig. 5: Qualitative statistical analysis of the relevance of bibliographic sources cited by the patient-oriented GPT model

Fig. 5: Qualitative statistical analysis of the relevance of bibliographic sources cited by the patient-oriented GPT model

Fig. 6: Qualitative statistical analysis of the relevance of bibliographic sources cited by the hospital-oriented GPT model

Fig. 6: Qualitative statistical analysis of the relevance of bibliographic sources cited by the hospital-oriented GPT model

Evaluation of permanent biological impairment estimates

Assessment of permanent biological impairment estimates was performed exclusively in the five cases with confirmed healthcare liability. In these cases, the original medico-legal reports included impairment percentages provided independently by the Hospital’s consultants and the Patient’s experts. Detailed case-by-case scoring is provided in _Supplementary Material _9.

The patient-oriented GPT produced impairment estimates rated as highly concordant with human expert evaluations (score 3) in four out of five cases (80%), with one non-concordant estimate.

The hospital-oriented GPT produced one highly concordant estimate (20%), one concordant estimate (20%), and three non-concordant estimates (60%).

These findings indicate a higher level of agreement between the patient-oriented model and human expert assessments in cases involving established liability (Figs. 7 and 8).

Fig. 7: Qualitative statistical analysis of biological impairment percentages generated by the patient-oriented GPT model

Fig. 7: Qualitative statistical analysis of biological impairment percentages generated by the patient-oriented GPT model

Fig. 8: Qualitative statistical analysis of biological impairment percentages generated by the hospital-oriented GPT model

Fig. 8: Qualitative statistical analysis of biological impairment percentages generated by the hospital-oriented GPT model

Statistical findings

Across all ten cases, the paired comparison of argumentative coherence scores showed no overall difference between the patient‑ and hospital‑oriented models (Wilcoxon signed‑rank: n = 8 non‑tied pairs, p = 0.844, r₍rb₎ = −0.08). Stratification by liability status revealed the expected orientation effect: in non‑liability cases (cases 1–5), the hospital‑oriented model tended to score higher (Wilcoxon: n = 5, p = 0.438, r₍rb₎ = −0.40), whereas in liability cases (cases 6–10) the patient‑oriented model tended to score higher (Wilcoxon: n = 3, p = 0.500, r₍rb₎ = +0.67).

Inter‑model agreement on clinical synthesis (1–3) was substantial‑to‑almost perfect (Cohen’s κ = 0.804), confirming consistent reconstruction of the clinical course by both models. Conversely, agreement on impairment estimates in liability cases (five cases) was slight (Cohen’s κ = 0.167), reflecting the known sensitivity of quantitative impairment appraisal to the evaluative stance.

At the citation level (123 AI‑generated references parsed), the patient‑oriented model produced 17/56 fabricated or non‑verifiable citations versus 7/67 for the hospital‑oriented model; the difference was statistically significant (Fisher’s exact test: odds ratio = 3.74, p = 0.0065). In contrast, the ordinal relevance of citations (0–3) did not differ significantly between models (Mann–Whitney U = 1639, p = 0.195, r = 0.109; small effect), suggesting that while formal verifiability is uneven, topical pertinence is broadly comparable.