Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2509 - Safety Risk Analysis with Generative Large-Language Models

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 31
POSTER

Presenter(s)

Sharareh Koufigar, MS Headshot
Sharareh Koufigar, MS - University of Washington, Seattle, WA

S. Koufigar1, M. Nguyen2, J. Kang3, and E. C. Ford4; 1Department of Radiation Oncology, University of Washington/ Fred Hutchinson Cancer Center, Seattle, WA, 2Crozer-Chester Medical Center, Upland, PA, 3Department of Radiation Oncology, University of Washington/Fred Hutchinson Cancer Center, Seattle, WA, 4Department of Radiation Oncology, University of Washington and Fred Hutchinson Cancer Center, Seattle, WA

Purpose/Objective(s): Risk analysis often involves expert analysis of corpuses of narrative text using Failure Mode and Effects Analysis (FMEA) or other methods. As such, large language models (LLMs) have shown promise in automating complex clinical tasks, but their utility for standardized FMEA scoring has not been well characterized. This study evaluates whether an off-the-shelf commercial LLM can reliably replicate expert consensus of FMEA scores.

Materials/Methods: As input data we use 112 failure mode narrative texts related to photon/electron EBRT publicly available from AAPM Task Group-275, representing expert consensus scores for severity (S), occurrence (O) and detectability (D). Failure modes were randomly partitioned into a training set (n=54) and a validation set (n=58). FMEA scoring was performed using OpenAI GPT-5 accessed through an institutionally approved, HIPAA-compliant instance. Zero-shot (no training examples) and multi-shot prompting (with training set examples) combined with debiasing and/or emotion prompting strategies were compared. Agreement between LLM-assigned and expert consensus scores were assessed using weighted Cohen's kappa for S, O, and D individually and Spearman rank correlation analysis for Risk Priority Number (RPN).

Results: Multi-shot combined with differential and emotion prompting achieved the best overall performance, with substantial agreement for RPN (Spearman ?=0.66, p<0.001), moderate agreement for severity (?=0.58), and substantial agreement for occurrence (?=0.64). However, detectability agreement was more limited (?=0.48). Several prompting strategies produced small or negative kappa values, indicating that zero-shot prompting yields clinically unreliable safety scoring.

Conclusion: This work demonstrates promise for general-purpose LLMs in FMEA scoring for radiation oncology, though performance was modest at best. Severity and detectability showed the highest variability, with detectability being the most difficult component to estimate, reflecting its dependence on process-based reasoning grounded in institution-specific workflow knowledge. These findings suggest that task-specific fine-tuning may be necessary before LLMs can reliably support clinical risk assessment.