Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2506 - Agentic AI with Automated Self-Evaluation for Radiation Oncology Physics Procedure Retrieval and Guidance

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 24
POSTER

Presenter(s)

Seonghoo Kim, MS, BS - Northwestern University, Chicago, IL

S. Kim1, P. Botsford2, N. Murphy3, J. Wong3, P. Yadav3, Y. Balasanov1, M. Abazeed3, I. J. Das3, and P. Teo3; 1Machine Learning and Data Science, McCormick School of Engineering, Northwestern University, Evanston, IL, 2IS Digital Solutions, Northwestern Medicine, Chicago, IL, 3Department of Radiation Oncology, Feinberg School of Medicine, Northwestern University, Chicago, IL

Purpose/Objective(s):

We developed a retrieval-augmented language model (RAG) grounded in in-house radiation oncology physics and dosimetry procedures to generate source-attributed guidance for treatment and QA workflows, enabling physicists to rapidly reference institution-specific protocols when uncertainty arises. An LLM-as-a-judge assessed performance. We hypothesized that LLM-based self-evaluation would enable a reliable agentic AI to guide physics procedures.

Materials/Methods:

Using institution-specific physics procedure documents spanning chart checking, electron therapy, eye plaque, GammaTile, Gamma Knife, HDR/LDR brachytherapy, in-vivo dosimetry, patient-specific, machine and CT-Sim QA, radiation safety, TBI/TSET, and treatment planning, we constructed a RAG database with semantic chunking (95% similarity threshold). 197 clinically realistic queries were generated. In a zero-shot setting, three LLMs generated responses restricted to the RAG database. A separate LLM-as-a-judge scored responses for accuracy, completeness, and adherence (scale:1–5), with human scoring (5 physicists) to confirm concordance and support automated model selection.

Results:

LLM-assigned scores for accuracy, completeness, and adherence were 4.1 ± 1.4, 3.9 ± 1.4, and 4.1 ± 1.4 for Amazon Nova 2.0 Lite; 4.3 ± 1.1, 4.0 ± 1.2, and 4.2 ± 1.2 for Meta Llama 3.3 70B; and 4.6 ± 0.9, 4.4 ± 1.1, and 4.5 ± 1.1 for OpenAI GPT-4.1, respectively. Although GPT-4.1 achieved higher mean LLM-assigned scores, comparison with human evaluation revealed significantly higher completeness and adherence ratings by the LLM-as-a-judge (p = 0.04 and p = 0.032, respectively), whereas Nova 2.0 Lite and Llama 3.3 70B demonstrated concordance with human scoring (Table 1). Higher LLM-as-a-judge scores for GPT-4.1 may relate to its longer mean responses (874 vs. 593 characters for Llama 3.3) and possible prompt–model alignment effects. Performance declined for procedures with lengthy narratives or embedded images, likely due to context fragmentation and the current text-only retrieval framework, as multimodal retrieval remains under development.

Conclusion:

A retrieval-augmented language model grounded in institution-specific radiation oncology physics procedures demonstrated strong performance for generating protocol-concordant guidance. An integrated LLM-as-a-judge framework enabled scalable automated evaluation, with overall concordance to human scoring, supporting its potential utility for model selection and deployment in clinical physics workflows.

Table 1. LLM performance for radiation oncology physics procedure guidance (n =197 queries)

Nova 2.0 Lite

Llama 3.3 70B

GPT 4.1

Accuracy

Completeness

Adherence

Accuracy

Completeness

Adherence

Accuracy

Completeness

Adherence

LLM-as-a-judge

4.1 ± 1.4

3.9 ± 1.4

4.1 ± 1.4

4.2 ± 1.2

4.0 ± 1.2

4.2 ± 1.2

4.6 ± 0.9

4.4 ± 1.1

4.5 ± 1.1

Human

4.0 ± 1.5

3.9 ± 1.5

4.0 ± 1.5

4.2 ± 1.3

4.0 ± 1.3

4.2 ± 1.2

4.3 ± 1.2

4.2 ± 1.2

4.2 ± 1.2

p-value

0.40

0.95

0.68

0.57

0.75

1.00

0.01

0.04

0.03