Work overview

Section 02 of 06

Materials and methods

Generative artificial intelligence in forensic medicine: a pilot study on AI-simulated medico-legal reports in healthcare liability cases

Federica Ministeri, Massimiliano Esposito, Martina Francaviglia, Lucio Di Mauro, Grazia Giulia Pantè, Monica Salerno, Cristoforo Pomara, and Francesco Sessa · 2026

Contents

Section 02 of 06

  1. 01Introduction
  2. 02Materials and methods
  3. 03Results
  4. 04Discussion
  5. 05Conclusions
  6. 06Electronic Supplementary Material
Text size
Work overview

Section 2 of 6

Materials and methods

Federica Ministeri, Massimiliano Esposito, Martina Francaviglia, Lucio Di Mauro, Grazia Giulia Pantè, Monica Salerno, Cristoforo Pomara, and Francesco Sessa · about 5 minutes

Study design, setting, and case selection

This pilot study was conducted at the Institute of Legal Medicine of the University of Catania and was designed as a retrospective, qualitative, comparative analysis. A retrospective analysis of medico-legal litigation cases examined at the Institute between 2020 and 2025 was performed. The initial dataset consisted of 350 cases, which were subsequently assessed for eligibility according to the predefined selection criteria.

Cases were included if all of the following criteria were met:

Availability of a medico-legal report issued by the Institute of Legal Medicine on behalf of the involved healthcare institution;Availability of a medico-legal report prepared on behalf of the patient or a detailed compensation claim supported by medico-legal arguments;

Cases were excluded if:

The patient died during the healthcare episode (claims involving death, for which no estimation of permanent biological damage was applicable);Incomplete or unavailable clinical documentation related to the healthcare episode under dispute.

After application of these criteria, 10 cases were included in the final analysis. Five cases had been previously classified as non-liability cases by the Institute of Legal Medicine, resulting in dismissal of compensation claims, while the remaining five cases involved confirmed healthcare liability with an associated estimate of permanent biological impairment.

All documentation was fully anonymized prior to analysis.

Configuration of customized GPT models

Two customized GPT-based models were developed using the ChatGPT platform (OpenAI) [20] to simulate opposing medico-legal perspectives typically encountered in healthcare litigation:

A patient-oriented model (“AI in Defense of the Patient”);A hospital-oriented model (“AI in Defense of the Hospital”).

In this study, the term GPT refers to a generative large language model capable of processing and generating natural language. The models were configured through prompt-engineering and document injection, by uploading curated medico-legal reference materials and model-specific instructional files.

Both models were provided with the same core set of authoritative medico-legal sources addressing healthcare liability, causal assessment, and evaluation of biological damage (Supplementary Material 1). In addition, each model received a dedicated set of written instructions, designed to define its evaluative stance and guide the structure and content of the generated medico-legal reports (Supplementary Material 2-3-4).

The models were used in their default configuration without parameter tuning (e.g., temperature or top-p), and no fine-tuning or retraining procedures were applied. The outputs were entirely determined by the provided documentation and predefined instructional prompts.

The patient-oriented model was supplied with instructions explicitly directing it to critically analyze clinical documentation, identify potential deviations from good clinical practice, and support hypotheses of healthcare liability when scientifically plausible.

Conversely, the hospital-oriented model received instructions aimed at emphasizing adherence to clinical guidelines and standards of care, contextualizing adverse events, and excluding or mitigating liability when justified by the available evidence.

Both models were instructed to produce formally structured medico-legal reports, including sections on clinical documentation, clinical course, medico-legal evaluation, conclusions, and bibliography. When generating bibliographic references, models were required to rely exclusively on clinical guidelines and scientific literature indexed in PubMed or Scopus and valid at the time of the healthcare event.

For each of the selected cases, identical anonymized clinical documentation (medical records, emergency department reports, discharge summaries, diagnostic findings, and follow-up data) was provided to both GPT models. No interactive prompts, clarifications, or corrective inputs were introduced during report generation.

Each model independently generated a simulated medico-legal report consistent with its predefined evaluative orientation. This approach was adopted to avoid operator-induced bias and to assess the autonomous performance of the configured models.

Evaluation framework

AI-generated reports were evaluated by a single medico-legal expert, using a structured, multi-level qualitative framework.

Comparison with human-authored medico-legal reports:

For each case, the medico-legal reports generated by the artificial intelligence models were directly compared with the corresponding reports authored by human medico-legal experts. Two qualitative scores were assigned:

Argumentative coherence and concordance (1–10), assessing alignment with the conclusions and reasoning of human experts;Quality of clinical synthesis (1–3), evaluating clarity, completeness, and accuracy in reconstructing the clinical course.

2.Evaluation of bibliographic references:

Each scientific reference generated by the models was assessed according to a predefined grid evaluating:

Existence of the cited article;Correctness of authorship, journal, and publication year;Validity of the DOI or bibliographic link;Relevance to the specific clinical case (scored on a 0–3 scale).

Fabricated or non-existent references were assigned a score of zero across all categories.

3.Evaluation of permanent biological impairment estimates:

In cases with confirmed healthcare liability, AI-generated estimates of permanent biological impairment were compared with those reported in human medico-legal evaluations. Concordance was scored on a 0–3 scale (non-concordant to highly concordant).

This evaluation framework enabled a comprehensive assessment of the coherence, reliability, and medico-legal consistency of AI-generated outputs.

The evaluation framework was specifically designed for this exploratory pilot study and was not previously validated. Therefore, the scoring system should be interpreted as an exploratory qualitative tool rather than a standardized measurement instrument.

Statistical analysis

Given the small sample size and the ordinal nature of the scoring scales, we performed an exploratory non-parametric analysis. Paired comparisons of argumentative coherence scores between the hospital-oriented and patient-oriented GPT models were assessed using the Wilcoxon signed-rank test overall and stratified by liability status (cases without confirmed malpractice vs. cases with confirmed malpractice). Effect sizes for paired comparisons were reported as rank-biserial correlation (r_rb), computed from the signed-rank decomposition.

Inter-model agreement on clinical synthesis scores (1–3) and on impairment estimates (0–3, liability cases only) was quantified using Cohen’s κ with unweighted categories. Fabricated or non-verifiable references were operationally defined as citation entries scoring 0 simultaneously on all four formal dimensions (authorship, journal, year, DOI/link). The proportion of fabricated references was compared between models using Fisher’s exact test on 2 × 2 contingency tables. To compare the ordinal distributions of citation relevance (0–3) between models, we applied a Mann–Whitney U test (noting that, for two groups, it is equivalent to a Jonckheere–Terpstra trend test). A two-sided α = 0.05 was adopted, and all analyses were interpreted as exploratory.