Section 1 of 6
Introduction
Federica Ministeri, Massimiliano Esposito, Martina Francaviglia, Lucio Di Mauro, Grazia Giulia Pantè, Monica Salerno, Cristoforo Pomara, and Francesco Sessa · about 2 minutes
Healthcare liability litigation has progressively increased over recent decades, reflecting growing patient expectations, rising clinical complexity, and enhanced judicial scrutiny of adverse healthcare events [1–4]. Within this context, medico-legal experts play a pivotal role at the intersection between medicine and law, being required to assess the appropriateness of healthcare conduct, determine causal relationships between alleged errors and harm, and quantify biological damage with scientific rigor and impartiality [5]. Despite the centrality of this function, medico-legal evaluations inherently involve interpretative judgment, often resulting in substantial variability between opposing expert opinions [6–8].
In response to this complexity, modern medico-legal practice has increasingly emphasized evidence-based reasoning, methodological transparency, and explicit reference to clinical guidelines and scientific literature valid at the time of the alleged event [9]. In Italy, this approach has been formally reinforced by Law no. 24/2017 (Gelli–Bianco Law), which strengthened the role of clinical guidelines and good practices in healthcare liability assessments and promoted structured risk management strategies [10]. Similar trends are observed across other legal systems, where courts and healthcare institutions increasingly expect medico-legal opinions to be grounded in verifiable scientific evidence rather than subjective expert interpretation alone [11–13].
Against this backdrop, variability in medico-legal reasoning remains a critical issue, particularly in adversarial contexts such as healthcare litigation. Differences in interpretative frameworks, selective use of literature, and divergent weighting of clinical facts can lead to markedly different conclusions even when experts analyze the same documentation. This variability represents not only a methodological challenge but also a potential source of uncertainty for judicial decision-making and institutional risk management [6, 14, 15].
In parallel, artificial intelligence (AI), and specifically large language models (LLMs) based on generative architectures, has rapidly expanded its presence in healthcare, demonstrating the ability to analyze complex textual data, synthesize information, and generate structured narratives [16–18]. While these systems have been investigated in several clinical domains, including documentation support, diagnostic assistance, and risk prediction, their application within forensic medicine and medico-legal consultancy remains largely unexplored. This gap is particularly relevant given that medico-legal practice is primarily based on the analysis of clinical documentation and on the evaluation of relevant scientific literature and clinical guidelines [19].
To date, no experimental studies have evaluated whether generative AI systems can simulate medico-legal reasoning by analyzing real clinical documentation and producing structured medico-legal reports comparable to those authored by human experts. Moreover, the potential of AI to reproduce divergent argumentative perspectives, such as those typically adopted by experts acting on behalf of patients or healthcare institutions, has not been systematically investigated.
The present study addresses this gap by exploring the use of two customized GPT-based models configured to simulate opposing medico-legal perspectives in real cases of alleged healthcare liability. By comparing AI-generated reports with human-authored medico-legal evaluations, this study aims to assess the coherence, reliability, and limitations of AI-simulated medico-legal reasoning, and to explore its potential role as a preliminary support tool in contemporary forensic practice, rather than as a substitute for human expertise.