Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2626 - RadOnc_Benchmark: Mapping the Jagged Knowledge Frontier of LLMs in Radiation Oncology Using Failure-Focused, Epoch-Based Benchmarking

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 21
POSTER

Presenter(s)

Nikhil Thaker, MD, MBA, MHA - Capital Health Medical Center Hopewell, Pennington, NJ

N. G. Thaker1, N. Redjal1, T. J. Royce2, J. Varughese1, R. Kraft1, V. Subbiah3, A. Loaiza-Bonilla4, and T. N. Showalter5; 1Capital Health, Pennington, NJ, 2The University of North Carolina at Chapel Hill, Chapel Hill, NC, 3Stanford Cancer Institute, Palo Alto, CA, 4St Luke's University Health Network, Allentown, PA, 5University of Virginia, Charlottesville, VA

Purpose/Objective(s): LLM performance on medical exams and clinical scenarios is often reported as aggregate accuracy, masking safety-critical failure regions. We developed RadOnc_Benchmark, an epoch-based evaluation framework to (1) isolate persistent “hard failures” in stable archival knowledge and (2) stress-test robustness as knowledge shifts from canonical standards to frontier evidence. We hypothesized that oncology knowledge is not a smooth overlap between models and clinical expertise, but a jagged boundary with unpredictable failure clusters—precisely where retrieval and context engineering should be targeted.

Materials/Methods: Three question sets were evaluated under strict zero-shot conditions (no retrieval, no fine-tuning). Legacy: 1,483 ACR TXIT questions (2016–2021) were used to derive a Hard_Set of 227 persistently missed items. Canonical: 200 challenging questions curated from ASCO/ASTRO-supported educational resources. Frontier: 200 questions derived from practice-changing oncology studies from the prior 12 months (conference highlights, abstracts/manuscripts, ASCO/ASTRO digests). We benchmarked proprietary and open-weight models spanning multiple scales and “reasoning effort” configurations. Binomial confidence intervals were computed for key proportions; between-set performance shifts were assessed using conservative two-proportion comparisons.

Results: Performance depended strongly on knowledge epoch. On the Legacy Hard_Set, accuracy ranged 25.6–54.6%, with best performance 54.6% (124/227; 95% CI 48.1–61.0%), indicating a large, reproducible failure region even among the latest frontier models. On Canonical knowledge, top models approached ceiling performance (94–94.5%; e.g., 189/200, 95% CI 91.4–97.6%), while smaller/open-weight models remained heterogeneous. On Frontier knowledge, best accuracy plateaued at 86.0% (172/200; 95% CI 81.2–90.8%); exemplar models showed significant drops from canonical performance (e.g., 94.5%?84.0%); locally run open weight models mostly scored between 72-78%. Increasing “reasoning effort” produced modest gains at steep cost increases (e.g., 82.5% at $0.05 vs 86.0% at $0.36).

Conclusion: Oncology LLM competence is jagged and epoch-dependent: near-ceiling canonical performance coexists with large, persistent failure clusters in legacy hard cases and measurable brittleness on frontier evidence. This supports a shift from success-based benchmarks to failure-focused, nuanced evaluation to guide system design. Hard-set and frontier failures define where RAG, context engineering, and agentic verification should be applied to improve clinical reliability rather than merely inflating average scores.