2552 - A Comparative Evaluation of Treatment Management Recommendations for Gynecologic Malignancies by Four Large Language Models
Presenter(s)
N. Mladkova1, T. Y. Andraos2, E. Healy3, T. L. Smith4, A. Yaney5, and D. Fabian6; 1Westchester Medical Center, New York Medical College, Valhalla, NY, 2Department of Radiation Oncology, The Ohio State University Wexner Medical Center, Columbus, OH, 3Department of Radiation Oncology, University of California Irvine, Orange, CA, 4Memorial Cancer Institute, Hollywood, FL, 5Mount Carmel Health System, Grove City, OH, 6University of Kentucky, Department of Radiation Medicine, Lexington, KY
Purpose/Objective(s):
Large language models (LLMs) represent a significant advancement in the field of artificial intelligence capable of performing various language-processing tasks, yet their ability to generate clinically valid answers for treatment recommendations in gynecologic malignancies including radiation therapy management remains unexplored. This study evaluated the quality of LLM-generated treatment recommendations for gynecologic tumors across four commercially available models.Materials/Methods:
Six board-certified radiation oncologists evaluated LLM-generated therapeutic recommendations for a set of five de-identified gynecologic malignancy cases. Each treatment domain (EBRT technique, brachytherapy technique, EBRT dose, brachytherapy dose, fractionation, systemic treatment, target volume delineation, organ-at-risk identification, organ-at-risk dose constraints) was scored on a three-point scale (0 = inappropriate, 1 = partially appropriate, 2 = appropriate), and overall treatment grade on a ten-point scale (0 = clinically dangerous, 10 = perfect). A repeated-measures ANOVA with Huynh-Feldt correction was used to test for differences among models, with patient case as a blocking factor.Results:
OpenAI GPT-5.2 achieved the highest mean aggregate score (1.714 +/- 0.171), followed by Google Gemini 3.1 Pro (1.556 +/- 0.109), Claude Sonnet 4.6 (1.527 +/- 0.213), and DeepSeek-R1 (1.263 +/- 0.331). The overall effect of the model was statistically significant (F(3,12) = 5.21, p = .032). No individual pairwise comparison reached significance after Bonferroni correction. GPT-5.2, Gemini 3.1 Pro, and Sonnet 4.6 performed comparably, while DeepSeek-R1 scored consistently lower with notably greater variability across cases. Significant differences were noted across treatment domains (p<0.001) with Target Volumes and OAR Constraints receiving the lowest scores. Notably, no recommendation generated by GPT-5.2 was rated as clinically dangerous, while at least two plans by other models achieved an overall treatment grade of 3 or lower.Conclusion:
The currently available LLMs evaluated showed a variable capability in generating clinically appropriate medical recommendation, with none achieving consistent scores across all domains tested. GPT-5.2 produced the highest-quality recommendations overall, while DeepSeek-R1 demonstrated significant deficiencies. These findings suggest that leading LLMs may have the potential to serve as clinical decision-support or educational tools, and emphasize the necessity of human expert oversight in LLM-generated medical content.