2452 - Performance of GPT-4.1 on Clinical Radiation Oncology Board-Style Questions: Accuracy, Reasoning Quality and Error Patterns
Presenter(s)
V. Fahmy1, and J. S. Welsh2; 1Department of Radiation Oncology, Stritch School of Medicine, Loyola University Chicago, Maywood, IL, 2ProTom International, Flower Mound, TX
Purpose/Objective(s): Large language models (LLMs) are increasingly explored for medical education and clinical decision support, yet their performance in specialty-specific domains such as radiation oncology remains incompletely characterized. We evaluated GPT-4.1 on clinical radiation oncology board-style questions, assessing accuracy, reasoning quality, and error patterns. We hypothesized that GPT-4.1 would achieve substantial accuracy with higher reasoning scores for correct versus incorrect responses, and that errors would primarily reflect guideline misapplication or disease-specific nuances.
Materials/Methods: One hundred multiple-choice clinical radiation oncology questions across five disease sites (breast, prostate, CNS, gastrointestinal, benign/palliative) were curated from publicly available American College of Radiology board-preparation materials (2016–2025). Each question was input to GPT-4.1 to generate an answer and explanation. Two board-certified radiation oncologists independently graded correctness (correct/incorrect) and reasoning quality (0–2 scale), with discrepancies resolved by consensus. Incorrect responses were categorized by error type. Analyses included overall accuracy with 95% confidence intervals, chi-square testing for site differences, t-tests for reasoning scores, and Cohen’s kappa for inter-rater reliability.
Results: Overall accuracy was 71% (95% CI 61–79%), highest in breast (80%) and lowest in benign/palliative conditions (60%), without statistically significant differences across disease sites (?²(4)=5.7, p=0.22). Mean reasoning score was 1.52 ± 0.49, significantly higher for correct versus incorrect responses (1.81 vs. 0.74, p<0.001). Inter-rater reliability was excellent (?=0.88). Among 29 incorrect responses, errors were primarily due to guideline misapplication (45%) and disease-specific nuance (31%). Reasoning was coherent in 83% of all responses, including 62% of incorrect answers.
Conclusion: GPT-4.1 demonstrates substantial accuracy and high-quality reasoning on radiation oncology board-style questions, supporting its potential role in AI-assisted medical education. However, frequent guideline misapplication and lower accuracy in benign/palliative questions highlight the need for human oversight. These findings provide the first systematic assessment of LLM performance in specialty-specific training and inform future integration of AI tools in clinical education.
Key figure: Accuracy of GPT-4.1 on Radiation Oncology Board-Style Questions by Disease Site| Disease Site | Accuracy (%) | Reasoning Score (Mean ± SD) |
|---|---|---|
| Breast | 80 | 1.81 |
| Prostate | 72 | 1.52 |
| CNS | 70 | 1.48 |
| GI | 65 | 1.45 |
| Benign/Palliative | 60 | 1.32 |
| Overall | 71 | 1.52 |