Main Session
Sep 29
SS 34 - The Automated Clinic: AI-Powered Contouring, Monitoring, and Data Extraction

291 - Automated Identification and CTCAE Grading of Radiation-Induced Pneumonitis Using Large Language Models

03:05pm - 03:15pm ET
Room 257

Presenter(s)

Junyu Li, MD Headshot
Junyu Li, MD - Mayo Clinic Arizona, Phonix, AZ

J. Li1,2, Y. Ding1, D. A. S. Toesca1, N. Y. Yu1, R. Tao1, W. G. Rule1, J. B. Ashman1, T. T. W. Sio1, and W. Liu1; 1Department of Radiation Oncology, Mayo Clinic, Phoenix, AZ, 2Jiangxi Cancer hospital, NANCHANG, China

Purpose/Objective(s):

Radiation-induced pneumonitis (RIP) surveillance in thoracic oncology is hindered by fragmented electronic medical record documentation and complex medical comorbidities. Distinguishing RIP from COPD exacerbations or infectious processes requires contextual interpretation across longitudinal unstructured clinical narratives. Manual chart abstraction is labor-intensive and prone to inter-rater variability, limiting scalability for toxicity research. We developed a Large Language Model (LLM)–based pipeline to automatically identify and grade RIP from routine clinical documentations.

Materials/Methods:

We retrospectively analyzed 451 patients with thoracic malignancies (163 with lung primaries; 288 with esophageal primaries) treated with radiation therapy (RT) between 2009–2024. Among them, 101 patients developed RIP (44 grade 1; 57 grade =2). An in-house physician-informed prompt guided LLM pipeline developed to process longitudinal clinical notes within a monitoring window spanning from RT initiation until 180 days post RT completion. The prompt prioritized explicit diagnostic labels, imaging findings, and steroid usage as a severity marker for RIP. A predefined document-importance hierarchy was applied to identify RIP events and assign CTCAE grades, with outputs compared against physician-reviewed reference standards.

Results:

The LLM pipeline identified RIP with high fidelity, achieving an accuracy of 95.57%, a sensitivity of 96.04%, specificity of 95.43%, and a F1-score of 0.9065. Weighted Kappa for grading agreement was 0.8987 (95% CI: 0.8576–0.9324), indicating near-perfect agreement with physician review. Subgroup analysis demonstrated robust performance in both esophageal (k=0.9002, accuracy=97.22%) and lung cancer (k=0.8881, accuracy=92.64%) cohorts. For grade =2 RIP events, the model achieved an accuracy of 96.90%, a sensitivity of 91.23%, and a F1-score of 0.8814. Notably, sensitivity reached 100% for grade =2 RIP events in the esophageal subgroup, ensuring no high-grade toxicities were missed.

Conclusion:

Our LLM-based pipeline accurately identified and graded RIP from electronic medical records, enabling scalable and standardized toxicity surveillance for thoracic radiotherapy.