Main Session
Sep 30
QP 35 - Digital Allies: AI Tools Built for the Real Clinical Environment

1205 - Guideline Grounded Agentic Retrieval System Performance among Fourteen Radiation Oncology Disease Sites

08:05am - 08:10am ET
Room 160

Presenter(s)

Kishan Patel, MD Headshot
Kishan Patel, MD - McGaw Medical Center of Northwestern University, Chicago, IL

K. Patel1, R. Gong2, M. He2, S. Zhong2, S. Kim2, Y. Liu1, Y. Balasanov3, M. Abazeed1, and T. T. Peng1; 1Department of Radiation Oncology, Feinberg School of Medicine, Northwestern University, Chicago, IL, 2Master of Science in Machine Learning and Data Science, McCormick School of Engineering, Northwestern University, Evanston, IL, 3Machine Learning and Data Science, McCormick School of Engineering, Northwestern University, Evanston, IL

Purpose/Objective(s):

Clinical decision-making in radiation oncology requires precise interpretation of complex treatment guidelines, yet general large language models (LLMs) are not optimized for specialty reasoning and may generate inaccurate or non–guideline-concordant recommendations. To address this gap, an agentic retrieval-augmented generation system grounded in radiation oncology guidelines was developed. We hypothesized that a domain-anchored architecture would achieve reliable performance with robust concordance between LLM-based and physician evaluations, supporting scalable validation and safe clinical integration.

Materials/Methods:

Guidelines from 14 disease sites were extracted from curated clinical reference documents and converted into tiered data. Text was then segmented based on document structure (e.g., section titles/headings) and divided into sentence units =300 tokens to preserve context, indexed in a cloud database, and embedded using a transformer-based model with cosine-similarity retrieval. A seven-stage pipeline included PHI guardrails, intent classification, task decomposition, dense retrieval, citation-grounded synthesis (70B model), LLM-based self-assessment, and output validation. Performance was evaluated using 204 queries spanning fourteen disease domains. Two radiation oncologists rated responses for accuracy, completeness, adherence (1–5), and potential harm (0–4). LLM-based grading was compared with physician ratings using paired Wilcoxon tests.

Results:

Physician scores were high: accuracy 4.82±0.58, completeness 4.81±0.50, adherence 4.93±0.38. Potential harm occurred in 6.9% of responses. LLM-based scoring assigned higher scores for accuracy (4.97±0.21 vs 4.82±0.58, p=7.9×10?5) and adherence (4.99±0.10 vs 4.93±0.38, p=0.0097) but lower for completeness (4.44±0.60 vs 4.81±0.50, p=1.17×10?¹7). Completeness differences occurred in 12/14 sites (86%, p<0.05), whereas accuracy and adherence showed no site-specific variation. Agreement was high for accuracy (90.7%) and adherence (96.1%) but lower for completeness (60.3%). Rank correlations were moderate across domains (?˜0.40–0.50, all p<10?8).

Conclusion:

A guideline-grounded agentic retrieval architecture achieved near-expert performance across 14 radiation oncology domains with low harm rates, establishing the feasibility of clinically reliable language model decision support. LLM–physician grading differences were systematic and primarily reflected stricter completeness thresholds, not performance instability. Domain-anchored retrieval frameworks generate accurate, guideline-concordant responses and facilitate scalable validation and clinical integration, enhancing consistency, safety, and efficiency in decision-making. This represents the first multi-disease-site demonstration of near-expert guideline-grounded language model performance in radiation oncology, setting a new benchmark for clinical AI.