Section 4 of 6
Discussion
Federica Ministeri, Massimiliano Esposito, Martina Francaviglia, Lucio Di Mauro, Grazia Giulia Pantè, Monica Salerno, Cristoforo Pomara, and Francesco Sessa · about 5 minutes
This pilot study represents an initial empirical assessment of the ability of generative artificial intelligence to simulate medico-legal reasoning in real cases of alleged healthcare liability. The findings show that GPT-based models, when configured with divergent evaluative orientations, are capable of producing structured medico-legal reports from identical clinical documentation, with outputs that vary predictably according to the assigned perspective. The objective was to evaluate whether generative artificial intelligence models, configured to adopt predefined medico-legal perspectives, could analyze real clinical documentation and generate simulated medico-legal reports comparable to those produced by human forensic experts in the context of healthcare liability.
The overall comparison of argumentative coherence revealed no statistically significant difference between the patient-oriented and hospital-oriented models (Wilcoxon signed-rank test, p = 0.844). This finding suggests that, when considered globally, both models achieve a comparable level of alignment with human-authored medico-legal reports.
When cases were stratified by liability status, a directionally consistent orientation effect emerged. In non-liability cases, the hospital-oriented model tended to achieve higher coherence scores, whereas in liability-confirmed cases the patient-oriented model tended to perform better. Although these differences did not reach statistical significance (p = 0.438 and p = 0.500, respectively), effect size estimates indicated a moderate-to-large directional trend, particularly in liability cases (r₍rb₎ = +0.67). This pattern is consistent with the predefined evaluative conditioning of the models and demonstrates that role assignment systematically influences AI-generated medico-legal reasoning, even in the absence of statistically significant differences due to limited sample size [21, 22].
Inter-model agreement analyses further delineate the functional domains in which generative AI performs more reliably. Agreement between models was substantial to almost perfect for clinical synthesis (Cohen’s κ = 0.804), confirming that both systems consistently reconstructed the clinical course from identical documentation. This finding confirms that generative AI is particularly effective in organizing heterogeneous clinical records into coherent narratives, a task that represents a substantial component of routine medico-legal activity.
However, this strength must be interpreted cautiously. In cases characterized by complex clinical trajectories, prolonged timelines, or subtle medico-legal nuances, the models occasionally failed to capture elements that would be considered critical by human experts, such as borderline deviations from best practice, ambiguous temporal relationships, or context-dependent clinical decisions. This limitation reflects the models’ reliance on textual pattern recognition rather than on conceptual or causal understanding, and it reinforces the necessity of expert interpretation when AI-generated summaries are used in medico-legal contexts.
Conversely, agreement on permanent biological impairment estimates was slight (Cohen’s κ = 0.167), highlighting a critical limitation. The estimation of impairment is inherently evaluative and normative, requiring integration of clinical findings, causal reasoning, functional assessment, and medico-legal standards. The low concordance observed confirms that AI-generated quantitative outputs in this domain are highly sensitive to evaluative orientation and cannot be considered reliable without expert oversight.
The statistical analysis of bibliographic outputs revealed one of the most clinically relevant findings of the study. The patient-oriented model generated a significantly higher proportion of fabricated or non-verifiable citations compared with the hospital-oriented model (OR = 3.74; p = 0.0065). This statistically significant difference confirms that bibliographic reliability is not uniform across role-conditioned models and may be influenced by the argumentative stance imposed during configuration. In contrast, no significant difference was observed in citation relevance between models (Mann–Whitney U test, p = 0.195), indicating that topical pertinence and formal verifiability represent distinct dimensions of bibliographic quality.
From a medico-legal perspective, this distinction is crucial. While AI systems may be capable of identifying thematically appropriate literature, the generation of inaccurate or non-existent references poses a substantial risk in forensic contexts, where evidentiary reliability and traceability are mandatory. In judicial contexts, the validity, traceability, and verifiability of scientific references are essential prerequisites for evidentiary reliability. The generation of non-existent or inaccurate citations by AI systems represents a critical limitation, as it may compromise the credibility of medico-legal reports and potentially affect judicial decision-making. In adversarial legal settings, where expert reports are subject to scrutiny by opposing parties and courts, even a single fabricated reference may undermine the overall reliability of the document. Therefore, strict human validation of all AI-generated bibliographic content is mandatory [14, 15, 23, 24].
Taken together, the results of this study indicate that generative AI may have a role in selected preparatory phases of medico-legal activity, such as the organization of clinical documentation, the drafting of structured summaries, and the simulation of alternative argumentative perspectives. However, the limitations observed, particularly with respect to interpretative depth of clinical data and bibliographic reliability, preclude autonomous use in medico-legal decision-making [22, 25, 26].
The integration of generative AI into forensic practice should be governed by clear frameworks emphasizing transparency, traceability, and mandatory human oversight, ensuring that technological support enhances efficiency without undermining professional responsibility [27, 28].
These findings also raise important ethical and legal considerations, particularly regarding responsibility, accountability, and transparency in AI-assisted medico-legal evaluations. The use of non-transparent generative systems in forensic contexts may challenge the requirement for explainability in legal proceedings, highlighting the need for clear regulatory frameworks.
Strengths and Limitations
As a pilot investigation, this study has several limitations. The sample size was small, reflecting the exploratory nature of the project, and included only cases with full documentation and clear expert reports. Additionally, the study relied on a single AI platform and a specific set of customized instructions. Different models or training materials could yield divergent outputs. All qualitative evaluations were performed by a single medico-legal expert, and no inter-rater reliability assessment was conducted. This represents a significant methodological limitation, particularly in medico-legal research, where interpretative variability between experts is well documented. Future studies should include multiple independent evaluators and formal agreement analysis to strengthen the robustness of the findings.
Finally, the models were not tested in dynamic or interactive contexts. However, this aspect also represents a methodological strength of the study, as the absence of any real-time interaction or iterative prompting ensured that the AI systems operated autonomously, relying exclusively on the preloaded instructions and source materials. This approach allowed for a more objective assessment of the models’ intrinsic reasoning and generative capabilities, independent of human guidance during the report generation process.