1125 - Directional Error Analysis of Large Language Model Triage for Radiation Oncology Incident Learning System Forms: Over-Triage vs. Under-Triage
Presenter(s)
S. Siddiqui1, S. J. Gardner2,3, C. Li3, B. M. Miller2, N. Baughan3, M. Dickinson4, A. M. Feldman1, A. J. Doemer1, B. Movsas1, and K. Thind1; 1Department of Radiation Oncology, Henry Ford Health, Detroit, MI, 2Henry Ford Health System, Detroit, MI, 3Henry Ford Health, Detroit, MI, 4Henry Ford Hospital, Detroit, MI, United States
Purpose/Objective(s): Incident learning systems (ILS) are recommended by ASTRO and AAPM as essential safety tools in radiation oncology, but manual triage remains a bottleneck at high-volume institutions. Prior approaches to automated ILS triage have been limited by small datasets and binary severity classification. We applied a large language model (LLM) to classify 96 ILS forms and evaluated the direction of its errors: over-triage (classifying events as higher risk than expert judges) versus under-triage (classifying as lower risk than expert judges).
Materials/Methods: Stratified spot-check sampling was used to select 96 of 500 ILS forms from a single academic radiation oncology department (2024-2025). Forms were classified by a 120B-parameter GPT-architecture LLM using zero-shot prompting for severity (Low/Medium/High) and dosimetric impact (Low/Medium/High). Three board-certified medical physicists independently rated agreement with each LLM prediction on a 5-point Likert scale. Over-triage rate was defined as the proportion of LLM High predictions rejected by expert majority (at least 2 of 3 Disagree/Strongly Disagree). Under-triage rate was the proportion of LLM Low predictions similarly rejected. Fisher’s exact test compared acceptance between High and Low predictions.
Results: Results are summarized in Table 1. Under-triage was rare (2.4-10.5% of Low predictions rejected) while over-triage was common (36-42% of High predictions rejected). For dosimetric impact, acceptance of Low (92.7%) versus High (33.3%) differed significantly (p=0.0001). Individual expert acceptance of High-dosimetric-impact predictions was similarly low (40-42%), compared to 80-95% for Low.
Conclusion: The LLM exhibited a safety-favorable error profile: when wrong, it predominantly over-triaged rather than under-triaged. Under-triage occurred in fewer than 11% of Low-risk predictions. This asymmetry supports deployment as a first-pass screen where Low classifications proceed with reduced oversight while Medium and High classifications receive expert review. The high over-triage rate, mirrored by low expert acceptance of the same High predictions, suggests high-acuity classification is inherently difficult for both human and AI reviewers and warrants standardized scoring rubrics.
Table 1. Directional Error Profile by LLM-Predicted Risk Level *=2/3 Agree/Strongly Agree. **=2/3 Disagree/Strongly Disagree. p = Fisher’s exact test, High vs Low| Dimension | LLM Level | n | Accept Rate* | Reject Rate** | Error Direction | p |
| Severity | Low | 38 | 76.3% | 10.5% | Under-triage | |
| Medium | 51 | 60.8% | 19.6% | — | ||
| High | 11 | 54.5% | 36.4% | Over-triage | 0.25 | |
| Dos. Impact | Low | 41 | 92.7% | 2.4% | Under-triage | |
| Medium | 47 | 48.9% | 27.7% | — | ||
| High | 12 | 33.3% | 41.7% | Over-triage | 0.0001 |