Main Session
Sep 29
PQA 05 - Physics

2994 - Evaluating GPT-4.1 Accuracy and Reasoning on Radiation Oncology Physics Board-Style Questions

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 32
POSTER

Presenter(s)

Veronia Fahmy, DO, BS Headshot
Veronia Fahmy, DO, BS - Loyola University Medical Center, Maywood, IL

V. Fahmy1, and J. S. Welsh2; 1Department of Radiation Oncology, Stritch School of Medicine, Loyola University Chicago, Maywood, IL, 2ProTom International, Flower Mound, TX

Purpose/Objective(s): Large language models (LLMs) are increasingly explored for medical education and decision support, but their performance on specialty-specific knowledge in radiation oncology physics is uncharacterized. We hypothesized that GPT-4.1 would demonstrate substantial accuracy and coherent reasoning on board-style physics questions, while errors would reveal domain-specific limitations.

Materials/Methods: We conducted a cross-sectional evaluation of GPT-4.1 using 100 multiple-choice questions across five physics domains—dosimetry & units, imaging & planning, beam physics, safety & QA, and brachytherapy—curated from publicly available RAPHEX-style practice exams. Each question was input to GPT-4.1, producing an answer and reasoning. Two board-certified radiation oncologists independently scored correctness (correct/incorrect) and reasoning (0–2: 0=incorrect/incoherent, 1=partially correct, 2=fully correct), resolving discrepancies by consensus. The primary endpoint was overall accuracy; secondary endpoints included reasoning score and error type distribution. Statistical analysis included 95% CIs, Chi-square tests for topic differences, t-tests for reasoning scores, and Cohen’s kappa for inter-rater reliability.

Results: Overall accuracy was 72% (95% CI 63–80%), highest in safety & QA (80%) and lowest in brachytherapy (65%). Mean reasoning score was 1.56 ± 0.48, higher for correct versus incorrect responses (1.83 vs 0.77, p<0.001). Inter-rater reliability was excellent (?=0.87). Among 28 incorrect answers, errors were primarily principle misapplication (39%), calculation/unit errors (32%), misinterpretation (21%), and other (8%), though reasoning remained coherent in 60% of incorrect responses. Accuracy and reasoning scores by topic are shown in the accompanying figure.

Conclusion: GPT-4.1 achieved 72% accuracy on radiation oncology physics board-style questions with largely coherent reasoning; however, this level of performance leaves substantial room for improvement. Errors—most commonly incorrect application of principles and calculation mistakes—highlight important limitations. These findings support LLMs as supplemental educational tools, not independent decision-makers, and provide a benchmark for future comparisons in AI-assisted radiation oncology physics education.

Key Figure: GPT-4.1 accuracy, reasoning scores, and error breakdown across 100 board-style radiation oncology physics questions.

Physics Topic

Questions (n) Accuracy (%) Mean Reasoning (Correct) Mean Reasoning (Incorrect)
Dosimetry & Units 20 75 1.85 0.78
Imaging & Planning 20 70 1.80 0.70
Beam Physics 20 68 1.78 0.72
Safety & QA 20 80 1.90 0.76
Brachytherapy 20 65 1.75 0.69
Overall 100 72 1.83 0.77