Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2655 - Large Language Model-Assisted Triage of Thoracic Radiation Oncology Referrals: A Pilot Validation Study

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 22
POSTER

Presenter(s)

David Wu, MD Headshot
David Wu, MD - Stanford Cancer Institute Palo Alto, Palo Alto, CA

D. J. Wu1, Z. B. White II1, J. S. Fernandes2, M. Diehn3, B. W. Loo Jr1, and M. F. Gensheimer1; 1Department of Radiation Oncology, Stanford University School of Medicine, Stanford, CA, 2University of Toronto, Toronto, ON, Canada, 3Stanford University School of Medicine, Stanford, CA

Purpose/Objective(s):

Timely triage of radiation oncology referrals is critical. At many institutions, triage depends on new patient coordinators (NPCs) who lack formal oncologic training to synthesize clinical data and assign priority, resulting in inefficient triage, potentially delaying care. We hypothesized a large language model (LLM) using a custom prompt and attending-designed triage rubric could generate recommendations comparable to radiation oncologists.

Materials/Methods:

Thirty-three consecutive thoracic radiation oncology referrals at a single academic cancer center were screened; 16 were selected to create a balanced set across three priority levels. Per-case clinical documentation (referral text, referring provider notes, and medical and/or surgical oncology notes) was input to a LLM (Gemini 2.5 Pro in HIPAA-approved enterprise environment) with a custom, iteratively refined prompt. The LLM generated an oncologic summary, triage recommendation (Urgent, Standard, Non-Urgent), and confidence band (High, Moderate, Low) per an attending-designed three-tier rubric. Two blinded reviewers (one fellow, one chief resident) independently triaged each case; discordant cases were adjudicated by a blinded attending radiation oncologist so that a gold-standard result was available for every case. Agreement was assessed by weighted Cohen's kappa. LLM summaries were compared to the clinically used NPC summaries using the 4Cs framework (Comprehensive, Concise, Coherent, Confabulation-free) on 5-point Likert scales via paired permutation test. This study was reviewed by our Institutional Review Board (IRB) and determined to be a quality improvement study not requiring IRB oversight.

Results:

Among 16 thoracic referrals (8 Non-Urgent, 6 Standard, 2 Urgent), the LLM achieved exact agreement with consensus in 75% of cases (?=0.65, 95% CI: 0.34–0.91), comparable to the fellow (?=0.72, exact agreement 81.2%) and chief resident (?=0.67, exact agreement 81.2%) individually; trainee inter-rater reliability was ?=0.39. All four LLM errors were off by one tier; among 15 high-confidence cases, the LLM was correct in 12 (2 under-triage, 1 over-triage), while the sole moderate-confidence case was also an error (under-triage). On 4Cs evaluation, LLM summaries outperformed NPC summaries on Comprehensiveness (4.44±0.61 vs 3.56±1.09, p=0.007), Conciseness (4.53±0.51 vs 2.84±0.80, p<0.001), and Coherence (4.22±0.47 vs 3.09±0.79, p<0.001); Confabulation-free scores were similar (4.72±0.47 vs 4.94±0.17, p=0.209).

Conclusion:

A custom-prompted LLM achieved substantial agreement with attending-adjudicated consensus (?=0.65), comparable to individual physician trainees (?=0.72, 0.67) and exceeded their inter-rater agreement (?=0.39), while significantly outperforming NPC summaries on 3/4 quality dimensions. These findings support broader, prospective evaluation of LLM-assisted triage as a component of new patient referral triage workflows in radiation oncology.