3645 - A Large-Language Model (LLM-based) Agent for Automatic Grading of Cancer Treatment-Related Toxicities from Multi-Encounter Clinical Notes
Presenter(s)
T. Upadhaya1, S. Armenia1, S. Chhetri2, J. Liu1, I. J. Chetty1, and K. M. Atkins1; 1Department of Radiation Oncology, Cedars-Sinai Medical Center, Los Angeles, CA, 2NepAl Applied Mathematics and Informatics Institute for Research (NAAMII), Kathmandu, Nepal
Purpose/Objective(s): Accurate and efficient assessment of cancer treatment-related side effects is essential for modeling event prediction, with the goal of optimizing the therapeutic ratio in routine treatment delivery—particularly in the era of multi-modal cancer therapies with novel intensified regimens where additive acute and late side effects are incompletely understood. Here we developed an LLM-based agent to automatically grade esophagitis and pneumonitis from multi-encounter clinical notes.
Materials/Methods: We utilized an institutional dataset (under IRB approval) for 2440 lung cancer patients, with 84324 clinical notes (average of 11 notes/patient). To establish a human reference standard, an expert radiation oncologist adjudicated CTCAE 5.0v grades for a patient subset: esophagitis (n=161) and pneumonitis (n=173). Cohen's Kappa (?), a statistic used to measure inter-rater reliability [range of -1 (high disagreement) to 1 (high agreement)], was used for the comparisons. A five-component orchestration framework using a locally deployed LLM (Llama 3.2 via Ollama) was developed consisting of: data processor, semantically driven note triage module for toxicity relevance, grading module, deterministic evidence aggregator, and a pipeline orchestrator to process multi-encounter patient records. Additionally, four LLM grading strategies were compared against each other: Retrieval-Augmented Generation (RAG) with a CTCAE vector store, few-shot learning (FSL, involving training using a small sample set), two-step reasoning, and CTCAE definition-guided prompting.
Results: Agreement between FSL and RAG strategies showed ?=0.19 (considered slight agreement) for esophagitis, and ?=0.29 (fair agreement) for pneumonitis. When compared against physician-adjudicated grades, FSL achieved ?=0.01 (esophagitis, slight agreement) and ?=0.41 (pneumonitis, moderate agreement); RAG achieved ?=0.09 (esophagitis, slight agreement) and ?=0.26 (pneumonitis, fair agreement). The indeterminant adjudication rate of the radiation oncologist due to insufficient information from the notes was 24 and 27%, for esophagitis and pneumonitis, respectively.
Conclusion: Our preliminary agentic LLM pipeline demonstrates feasibility for automated grading of cancer treatment-related toxicities from large-scale longitudinal clinical notes. This study highlights the promise and challenges, related to better training of the model, insufficient information in notes and the inclusion of larger samples. Future work will focus on expanding the comparison against physician-annotated ground truth, development of guidance for input of structured details in annotation of toxicities, and improving linkages to treatment-specific exposures (i.e., radiotherapy, cytotoxic chemotherapy, immunotherapy, targeted agents, etc).