Main Session
Sep 29
QP 21 - Innovating With Software to Drive Patient Safety & Quality

1125 - Directional Error Analysis of Large Language Model Triage for Radiation Oncology Incident Learning System Forms: Over-Triage vs. Under-Triage

04:10pm - 04:15pm ET
Room 256

Presenter(s)

Shayan Siddiqui, - Henry Ford Cancer Institute, Detroit, MI

S. Siddiqui1, S. J. Gardner2,3, C. Li3, B. M. Miller2, N. Baughan3, M. Dickinson4, A. M. Feldman1, A. J. Doemer1, B. Movsas1, and K. Thind1; 1Department of Radiation Oncology, Henry Ford Health, Detroit, MI, 2Henry Ford Health System, Detroit, MI, 3Henry Ford Health, Detroit, MI, 4Henry Ford Hospital, Detroit, MI, United States

Purpose/Objective(s): Incident learning systems (ILS) are recommended by ASTRO and AAPM as essential safety tools in radiation oncology, but manual triage remains a bottleneck at high-volume institutions. Prior approaches to automated ILS triage have been limited by small datasets and binary severity classification. We applied a large language model (LLM) to classify 96 ILS forms and evaluated the direction of its errors: over-triage (classifying events as higher risk than expert judges) versus under-triage (classifying as lower risk than expert judges).

Materials/Methods: Stratified spot-check sampling was used to select 96 of 500 ILS forms from a single academic radiation oncology department (2024-2025). Forms were classified by a 120B-parameter GPT-architecture LLM using zero-shot prompting for severity (Low/Medium/High) and dosimetric impact (Low/Medium/High). Three board-certified medical physicists independently rated agreement with each LLM prediction on a 5-point Likert scale. Over-triage rate was defined as the proportion of LLM High predictions rejected by expert majority (at least 2 of 3 Disagree/Strongly Disagree). Under-triage rate was the proportion of LLM Low predictions similarly rejected. Fisher’s exact test compared acceptance between High and Low predictions.

Results: Results are summarized in Table 1. Under-triage was rare (2.4-10.5% of Low predictions rejected) while over-triage was common (36-42% of High predictions rejected). For dosimetric impact, acceptance of Low (92.7%) versus High (33.3%) differed significantly (p=0.0001). Individual expert acceptance of High-dosimetric-impact predictions was similarly low (40-42%), compared to 80-95% for Low.

Conclusion: The LLM exhibited a safety-favorable error profile: when wrong, it predominantly over-triaged rather than under-triaged. Under-triage occurred in fewer than 11% of Low-risk predictions. This asymmetry supports deployment as a first-pass screen where Low classifications proceed with reduced oversight while Medium and High classifications receive expert review. The high over-triage rate, mirrored by low expert acceptance of the same High predictions, suggests high-acuity classification is inherently difficult for both human and AI reviewers and warrants standardized scoring rubrics.

Table 1. Directional Error Profile by LLM-Predicted Risk Level

*=2/3 Agree/Strongly Agree. **=2/3 Disagree/Strongly Disagree. p = Fisher’s exact test, High vs Low
Dimension

LLM Level

n

Accept Rate*

Reject Rate**

Error Direction

p

Severity

Low

38

76.3%

10.5%

Under-triage

Medium

51

60.8%

19.6%

—

High

11

54.5%

36.4%

Over-triage

0.25

Dos. Impact

Low

41

92.7%

2.4%

Under-triage

Medium

47

48.9%

27.7%

—

High

12

33.3%

41.7%

Over-triage

0.0001