Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2446 - Tuning the QA Dial: The Impact of Prompting and Ensemble Voting Thresholds on Autonomous LLM Peer Review in Breast Radiotherapy

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 15
POSTER

Presenter(s)

Claire Dvorak, - Orlando Health Cancer Institute, Orlando, FL

C. Dvorak1,2, and T. Dvorak1; 1Orlando Health Cancer Institute, Orlando, FL, 2Dartmouth College, Hannover, NH

Purpose/Objective(s): Large language models (LLMs) show promise as automated quality assurance (QA) tools for real-time peer review, yet safe clinic deployment requires clear operational thresholds. Because individual models exhibit distinct clinical biases, we hypothesized that epistemic uncertainty prompting and ensemble voting strategies would optimize accuracy while mitigating overtreatment bias and alert fatigue. We therefore evaluated their impact on foundational LLMs in breast radiotherapy.

Materials/Methods: Four frontier LLMs (GPT-5.2, Grok-4.1-Fast, Claude-Opus-4.5, Gemini-3.1-Pro) reviewed 100 consecutive unstructured breast cancer narratives from the EMR Oncologic History section for adjuvant radiotherapy (RT) indications and fractionation. Models were tested under “Closed World” (assume missing variables are pertinent negatives) versus “Open World” (acknowledge uncertainty) prompting. Performance was assessed individually and with ensemble voting (Unanimous 4-of-4 vs Majority 3-of-4) against a gold standard (Yes / Choice / Omission) established by a determination by a board-certified breast radiation oncologist.

Results:

Individual models demonstrated distinct clinical behaviors in varying thresholds for overtreatment and specificity, yielding moderate overall concordance to the gold standard (63%–73%). This was driven by a stark performance dichotomy: high sensitivity (~90%) for definitive RT (Yes), but poor recognition of de-escalation scenarios (Choice ~20%, Omission ~65%), with systematic overtreatment bias. Closed-World prompting exacerbated this bias; Open-World prompting reduced overconfidence, enabling Gemini-3.1-Pro to achieve 100% specificity for omitting RT and correctly identifying ~40% of shared-decision cases.

Ensemble analysis revealed a critical throughput-versus-precision trade-off. The Unanimous Open-World ensemble provided a precise “silent safety net” (83% overall accuracy) but reached consensus in only 59% of cases. The Majority ensemble expanded coverage to 84% of the clinic roster but dropped accuracy to 69%, incorrectly flagging 92% (22/24) of appropriate de-escalation cases as errors. All models also exhibited training-data lag: across 11 cases where 5-fraction ultra-hypofractionation (FAST-Forward) was used, its adoption was 0%, with universal recommendation of 15–28 fraction schedules.

Conclusion: Foundational LLMs may excel as safety nets for missed treatments but lack nuance for de-escalation or modern ultra-hypofractionation without specialized fine-tuning. For safe implementation, institutions can tune the QA dial: Majority ensemble maximizes case coverage to aggressively catch missed treatments, whereas Unanimous ensemble minimizes alert fatigue and protects appropriate de-escalation. Crucially, both settings require “Open-World” prompting to safely acknowledge clinical uncertainty, reserving “Closed-World” strictness for retrospective documentation audits.