Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2462 - Expert Validation of Large Language Model Triage Predictions for Radiation Oncology Incident Learning System Forms

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 30
POSTER

Presenter(s)

Stephen Gardner, MS - Henry Ford Health System, Detroit, MI

S. J. Gardner1, S. Siddiqui2, C. Li1, B. M. Miller3, N. Baughan1, M. Dickinson4, A. M. Feldman2, A. J. Doemer2, B. Movsas5,6, and K. Thind2,5; 1Henry Ford Health, Detroit, MI, 2Department of Radiation Oncology, Henry Ford Health, Detroit, MI, 3Henry Ford Health System, Detroit, MI, 4Henry Ford Hospital, Detroit, MI, United States, 5Medicine, Michigan State University, East Lansing, MI, 6Department of Radiation Oncology, Henry Ford Cancer, Detroit, MI

Purpose/Objective(s): Incident learning systems (ILS) are essential for quality improvement in radiation oncology, but manual triage of submitted forms is labor-intensive and subject to inter-reviewer variability. We hypothesized that a large language model (LLM) could generate triage classifications accepted by expert medical physicists at rates comparable to inter-expert agreement, enabling scalable initial screening of ILS forms.

Materials/Methods: A total of 96 ILS forms from a single academic radiation oncology department (2024-2025) were classified by 120B-parameter GPT-architecture LLM using zero-shot prompting across three dimensions: process step (8 categories), severity (Low/Medium/High), and dosimetric impact (Low/Medium/High). The forms were selected within the timeframe above to encompass a range of severity/dosimetric impact, with overall intent of selecting forms which contain relevant information for testing the effectiveness of LLM-based analysis. Three expert medical physicists independently evaluated each entry on 5-point Likert scale (Strongly Agree to Strongly Disagree). Inter-rater reliability was assessed using weighted Fleiss’ kappa (quadratic weights), Gwet’s AC2, Krippendorff’s alpha (ordinal), and intraclass correlation coefficients (ICC). Acceptance was defined as majority expert agreement (at least 2 of 3 rating Agree/Strongly Agree). Category-specific acceptance rates were calculated to identify where LLM triage succeeds and fails.

Results: Majority expert acceptance rates for LLM predictions were 81% (process step), 66% (severity), and 65% (dosimetric impact). Gwet’s AC2 indicated moderate concordance (0.513, 0.525, 0.600), while weighted Fleiss’ kappa values were lower (0.233, 0.143, 0.338) due to skewed distributions consistent with the kappa paradox. Performance was highly category-dependent. For process step, Simulation (100%) and Patient Setup (94%) achieved near-perfect acceptance, but Imaging (40%) and Treatment Planning (68%) fell well below acceptable thresholds. For dosimetric impact, Low predictions were accepted at 93%, but the LLM failed for higher-acuity classifications: Medium reached only 49% and High only 33%. Severity showed the weakest overall agreement, with Krippendorff’s alpha confidence interval crossing zero (-0.026 to 0.250), indicating reliability indistinguishable from chance.

Conclusion: LLM triage performed well for routine, low-acuity ILS classifications but failed to achieve acceptable expert agreement for high-acuity and severity predictions. These category-specific gaps preclude fully autonomous deployment and instead support a tiered strategy: LLM-only triage for well-validated categories, with mandatory expert review for severity, high dosimetric impact, and process steps with low acceptance rates.