2994 - Evaluating GPT-4.1 Accuracy and Reasoning on Radiation Oncology Physics Board-Style Questions
Presenter(s)
V. Fahmy1, and J. S. Welsh2; 1Department of Radiation Oncology, Stritch School of Medicine, Loyola University Chicago, Maywood, IL, 2ProTom International, Flower Mound, TX
Purpose/Objective(s): Large language models (LLMs) are increasingly explored for medical education and decision support, but their performance on specialty-specific knowledge in radiation oncology physics is uncharacterized. We hypothesized that GPT-4.1 would demonstrate substantial accuracy and coherent reasoning on board-style physics questions, while errors would reveal domain-specific limitations.
Materials/Methods: We conducted a cross-sectional evaluation of GPT-4.1 using 100 multiple-choice questions across five physics domains—dosimetry & units, imaging & planning, beam physics, safety & QA, and brachytherapy—curated from publicly available RAPHEX-style practice exams. Each question was input to GPT-4.1, producing an answer and reasoning. Two board-certified radiation oncologists independently scored correctness (correct/incorrect) and reasoning (0–2: 0=incorrect/incoherent, 1=partially correct, 2=fully correct), resolving discrepancies by consensus. The primary endpoint was overall accuracy; secondary endpoints included reasoning score and error type distribution. Statistical analysis included 95% CIs, Chi-square tests for topic differences, t-tests for reasoning scores, and Cohen’s kappa for inter-rater reliability.
Results: Overall accuracy was 72% (95% CI 63–80%), highest in safety & QA (80%) and lowest in brachytherapy (65%). Mean reasoning score was 1.56 ± 0.48, higher for correct versus incorrect responses (1.83 vs 0.77, p<0.001). Inter-rater reliability was excellent (?=0.87). Among 28 incorrect answers, errors were primarily principle misapplication (39%), calculation/unit errors (32%), misinterpretation (21%), and other (8%), though reasoning remained coherent in 60% of incorrect responses. Accuracy and reasoning scores by topic are shown in the accompanying figure.
Conclusion: GPT-4.1 achieved 72% accuracy on radiation oncology physics board-style questions with largely coherent reasoning; however, this level of performance leaves substantial room for improvement. Errors—most commonly incorrect application of principles and calculation mistakes—highlight important limitations. These findings support LLMs as supplemental educational tools, not independent decision-makers, and provide a benchmark for future comparisons in AI-assisted radiation oncology physics education.
Key Figure: GPT-4.1 accuracy, reasoning scores, and error breakdown across 100 board-style radiation oncology physics questions.| Physics Topic | Questions (n) | Accuracy (%) | Mean Reasoning (Correct) | Mean Reasoning (Incorrect) |
| Dosimetry & Units | 20 | 75 | 1.85 | 0.78 |
| Imaging & Planning | 20 | 70 | 1.80 | 0.70 |
| Beam Physics | 20 | 68 | 1.78 | 0.72 |
| Safety & QA | 20 | 80 | 1.90 | 0.76 |
| Brachytherapy | 20 | 65 | 1.75 | 0.69 |
| Overall | 100 | 72 | 1.83 | 0.77 |