Main Session
Sep 30
QP 35 - Digital Allies: AI Tools Built for the Real Clinical Environment

1209 - Reducing Documentation Burden in Radiation Oncology: Automated Treatment Summaries Using a Two-Stage Large Language Model System

08:25am - 08:30am ET
Room 160

Presenter(s)

Keldon Lin, MD, MS Headshot
Keldon Lin, MD, MS - Mayo Clinic College of Medicine and Science Rochester, Rochester, MN

K. K. Lin1, A. Y. K. Foong1, F. Mastroleo1, S. Seetamsetty1, M. Borras-Osorio1, S. Shiraishi1, J. Holmes2, W. Liu2, and M. R. Waddle1; 1Department of Radiation Oncology, Mayo Clinic, Rochester, MN, 2Department of Radiation Oncology, Mayo Clinic, Phoenix, AZ

Purpose/Objective(s): Radiation treatment summaries are critical for communicating a patient’s treatment course to other healthcare professionals and are mandated by ASTRO and the American College of Radiology for accreditation. However, their manual completion is time-consuming, redundant, and susceptible to variability in format and completeness. We evaluated the performance of a two-stage large-language model (LLM) system designed to automatically generate radiation oncology treatment summaries to meet documentation standards and reduce documentation burden.

Materials/Methods: Patients with pending or overdue treatment summary notes on the date of analysis were included. All clinical documents from the electronic health record including Care Everywhere notes (Epic) and radiation treatment details from the record and verify software (ARIA) were collected. First, Gemini 2.0 Flash concurrently processed these documents and condensed each into a four-sentence summary. Next, Gemini 2.5 Pro reviewed the flagged summaries and incorporated their timestamps to reconstruct the patient’s care timeline. The system then prompted Gemini 2.5 Pro with a formatted treatment summary template. The summaries were then assessed for completeness, accuracy, structure, conciseness, and adherence to ASTRO treatment summary guidelines using a Likert scale. Results were reported with descriptive statistics.

Results: A total of 62 generated radiation treatment summaries were analyzed from multiple disease sites and treatment types. Frequency analysis demonstrated near-maximum scores across most categories. In the Completeness category, 85.5% of cases scored maximum (5 points, mean=4.81±0.42), while the Structure and Conciseness categories exhibited higher scores (96.8% scoring 5 points each, means=4.97±0.18). The Accuracy category showed the greatest variability (mean=4.08±1.08). The ASTRO compliance score had a mean of 14.45±0.86 out of 15 possible points, where 59.7% achieved perfect scores (15 points) and 30.6% scored near perfect (14 points). In most cases where accuracy was low, it was due to errors in the total elapsed days, omission of concomitant chemotherapy, and missing treatment technique.

Conclusion: Our two-stage large language model system consistently generated treatment summaries that were complete, well-structured, concise, and highly compliant with ASTRO documentation standards. The greatest areas for improvement were (1) inclusion of treatment technique, (2) accuracy of concurrent chemotherapy, and (3) elapsed days of treatment. These represent a key opportunity for physician in-the-loop verification. Implementation of such LLM-based tools has the potential to reduce documentation burden, improve standardization, and enhance information sharing. Broader deployment and multi-institutional prospective validation are warranted to evaluate integration into clinical workflows.