Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2670 - Agentic Pipeline for Real-World Patient Queries in Radiation Oncology

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 22
POSTER

Presenter(s)

Abdul Zakkar, MD Headshot
Abdul Zakkar, MD - Northwestern Memorial Hospital, Chicago, IL

A. Zakkar1, P. D. Philip2, S. Kumar3, Y. Huang3, M. Morrison1, H. Nordhues4, Y. Balasanov5, P. Teo6, and M. Abazeed7; 1Department of Radiation Oncology, Northwestern University, Feinberg School of Medicine, Chicago, IL, 2Machine Learning & Data Science Program, Northwestern University, Evanston, IL, 3Northwestern University, Evanston, IL, 4Northwestern Memorial Hospital, Chicago, IL, 5Machine Learning and Data Science, McCormick School of Engineering, Northwestern University, Evanston, IL, 6Northwestern Medicine, Chicago, IL, 7Department of Radiation Oncology, Feinberg School of Medicine, Northwestern University, Chicago, IL

Purpose/Objective(s): Large language models (LLMs) have demonstrated progressively improved performance in retrieving and synthesizing medical knowledge, including examination-level proficiency in radiation oncology. Prior studies have also reported high accuracy, completeness, and conciseness of LLM-generated responses to standardized clinical questions derived from professional society resources (Yalamanchili et al., 2024). However, these evaluations have largely relied on curated or examination-style prompts rather than authentic patient inquiries. The performance and safety of LLMs when responding to real-world questions submitted by patients before, during, and after radiation therapy remain insufficiently characterized. In this study, we evaluate the safety and accuracy of agentic LLM responses to real-world radiation oncology patient queries.

Materials/Methods: We developed CancerRAG, an agentic retrieval-augmented generation (RAG) LLM designed to answer radiation oncology patient queries. The system retrieves context from a curated domain-specific knowledge base using modular LLM chains for knowledge extraction, relevance grading, and response synthesis, with outputs tailored to patient demographics. Its agentic architecture enables decomposition of complex queries and multi-step task management (e.g., appointment coordination). Responses were independently scored by a human expert for correctness, relevance, non-harmfulness, and coherence using a 5-point Likert scale, and the LLM performed parallel self-evaluation of the same four qualities.

Results: CancerRAG generated responses to 85 patient queries. Mean human expert scores (out of 5) were 4.01 for correctness, 4.48 for relevance, 4.72 for non-harmfulness, and 4.67 for coherence. Using the same Likert rubric, the LLM’s mean self-evaluation scores were 3.82, 3.57, 4.00, and 4.05, respectively. Queries with intent identified as “appointment” had responses which scored significantly higher by the human expert than those identified as “clinical” in correctness, non-harmfulness, and coherence.

Conclusion: CancerRAG delivered responses to radiation oncology patient queries with high levels of safety and accuracy, demonstrating particular strengths in non-harmfulness and coherence, consistent with prior LLM studies. Notably, the LLM’s self-evaluation scores were more conservative than those of the human expert. These findings support the prospective integration of agentic RAG pipelines into clinical workflows to augment patient care. Ongoing work includes additional expert evaluations and inter-rater agreement analyses.