2527 - Consistency of ChatGPT Responses for Stereotactic Radiosurgery Recommendations in AVM Management
Presenter(s)
C. Lin, and A. Saber; Department of Radiation Oncology, University of Nebraska Medical Center, Omaha, NE
Purpose/Objective(s): Artificial intelligence (AI) tools such as ChatGPT are increasingly used by clinicians and trainees to rapidly retrieve medical information, including guidance related to radiation therapy. However, the reliability and consistency of such outputs for clinical decision-making remain uncertain. To evaluate the consistency of ChatGPT-generated clinical recommendations for stereotactic radiosurgery (SRS) in the management of intracranial arteriovenous malformations (AVMs) when provided with the same query.
Materials/Methods: The identical prompt — “Please tell me how to treat AVM with SRS with detailed dose regimens for different situations” — was submitted to ChatGPT 20 separate times. After each response, the prior conversation and file were deleted before the next query was initiated. The 20 generated outputs were independently reviewed and qualitatively analyzed for concordance and variability in treatment principles and dose recommendations across different clinical scenarios, including small AVMs, large AVMs treated with staged SRS, brainstem AVMs, hypofractionated SRS, and the role of pre-SRS embolization.
Results: Strong concordance was observed in fundamental treatment principles, with all responses endorsing nidus-only targeting and delayed obliteration over 2–3 years following SRS. For small AVMs (=3 cm), 17 of 20 profiles recommended single-fraction SRS doses of 20–24 Gy. For large AVMs treated with volume-staged SRS, most profiles (12/20) recommended per-stage doses of 15–18 Gy, with smaller subsets favoring 14–16 Gy (4/20) or substantially lower doses of 10–12 Gy (1/20). In contrast, notable heterogeneity was observed for brainstem AVMs, with recommended single-fraction doses ranging from 12–14 Gy (5/20), 14–16 Gy (6/20), to 16–18 Gy (4/20), while several profiles did not specify numeric dose constraints. Additional variability was seen in recommendations regarding hypofractionated SRS and pre-SRS embolization, with inconsistent acknowledgment of inferior obliteration rates.
Conclusion: ChatGPT provides broadly consistent high-level principles for AVM SRS but demonstrates clinically meaningful variability in dose recommendations for critical locations and alternative fractionation strategies. These findings highlight that AI-generated outputs should not be used as a standalone source for clinical decision-making and underscore the importance of expert judgment and evidence-based guidelines when integrating AI tools into radiation oncology practice.