3272 - Clinical Validation of a Privacy-Preserving LLM for End-to-End CTCAE Toxicity Assessment in Prostate Cancer Radiotherapy
Presenter(s)
R. Khanmohammadi1, A. I. Ghanem2,3, A. R. Bhatnagar2, S. Siddiqui2, M. A. Elshaikh2, H. Bagher-Ebadian2,4, B. Movsas2,5, M. M. Ghassemi1, and K. Thind2,5; 1Department of Computer Science and Engineering, Michigan State University, East Lansing, MI, 2Department of Radiation Oncology, Henry Ford Health, Detroit, MI, 3Clinical Oncology Department, Faculty of Medicine, Alexandria University, Alexandria, Egypt, 4Department of Radiology, Michigan State University Medicine,, East Lansing, MI, 5Medicine, Michigan State University, East Lansing, MI
Purpose/Objective(s): Accurate adverse event (AE) assessment is critical in radiation oncology, yet manual Common Terminology Criteria for Adverse Events (CTCAE) grading is labor-intensive and inconsistent, limiting scalable prospective AE monitoring. We hypothesized that a compact, privacy-preserving large language model (LLM) could achieve clinically acceptable accuracy for automated end-to-end CTCAE extraction and grading across ten post-radiotherapy (RT) AEs, evaluated against expert annotations.
Materials/Methods: With IRB approval, we deployed an open-source 8B-parameter LLM in a two-stage zero-shot pipeline with chain-of-thought prompting: (1) AE extraction from clinical notes, followed by (2) CTCAE severity grading (Grades 1–3) for extracted positives. Inference ran locally to preserve patient privacy. Among 16,107 clinical notes of different specialities, from 217 prostate cancer patients treated with definitive external beam RT (74.6–79.2 Gy) along 5-7 years post-RT, 4,240 contained one or more RT-induced AEs. The pipeline was validated on 573 expert-annotated notes, yielding 1,799 pairs across 10 AEs (median 166 [150–185] per AE). Grading was assessed on correctly extracted positives (n=843) to isolate stage-specific performance. Primary metrics included per-grade F1 and combined Grade 2+3 F1 with bootstrap 95% CIs (1,000 resamples).
Results: Extraction achieved median F1=84% [70–92%] across AEs. Rectal bleeding and nocturia were best captured (F1=95%), while urgency (57%) and stricture (68%) proved most challenging due to term overlap and limited samples (Table 1). Grading achieved median G1=90%, G2=85%, G3=80%, and Grade 2+3 F1=88%, with best performance in urgency (95-96%), rectal bleeding (94-97%) and hematuria (85-95%). End-to-end G2+3 F1 across all 1,799 pairs was 78% [75–80%], with false high-grade predictions in 6% of pairs. Among grading errors, 96% were adjacent-grade, predominantly under-grading, indicating a conservative failure mode.
Conclusion: This validation demonstrates that a privacy-preserving LLM can reliably identify clinically actionable AEs, with grading errors that are predominantly conservative. While extraction remains the primary bottleneck, the pipeline’s safety profile supports deployment as a triage tool to automate toxicity surveillance from routine documentation, routing high-grade flags to clinician review saving time and resources. By surfacing structured AE data at scale, this approach may leverage awareness and timeliness of managing RT-related AE symptoms.
| AE | Extraction | Grading | ||||||
| F1 | N | G1 F1 | G2 F1 | G3 F1 | G2+3 F1 | |||
| Dysuria | 68 | 50 | 93 | 82 | 100 | 84 | ||
| Erectile Dysfunction | 91 | 126 | 91 | 85 | 80 | 88 | ||
| Hematuria | 92 | 83 | 92 | 85 | 91 | 95 | ||
| Incontinence | 88 | 76 | 86 | 77 | 56 | 87 | ||
| Nocturia | 94 | 132 | 86 | 82 | 100 | 83 | ||
| Rectal Bleeding | 95 | 96 | 95 | 94 | 81 | 97 | ||
| Stricture | 68 | 23 | 33 | 86 | 80 | 90 | ||
| Urgency | 56 | 68 | 95 | 96 | 0 | 96 | ||
| Urinary Frequency | 78 | 126 | 88 | 86 | 0 | 86 | ||
| Urinary Retention | 81 | 63 | 82 | 79 | 75 | 84 | ||
| OVERALL | Median | 84 | 80 | 90 | 85 | 80 | 88 | |
| IQR | 70-92 | 64-118 | 86-93 | 82-86 | 61-88 | 84-94 | ||