Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2516 - Automated Radiotherapy Data Extraction via a Verifiable LLM Agent-Orchestrated Pipeline: Integrating OIS SQL and TPS ESAPI Workflows

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 17
POSTER

Presenter(s)

Inbum Lee, PhD - Seoul National University Hospital, Seoul, SEO

I. Lee1,2, Y. Huh1,2, S. Kang1, J. I. Kim1,2, K. Kim3, and C. H. Choi1,2; 1Institute of Radiation Medicine, Seoul National University Medical Research Center, Seoul, Korea, Republic of (South), 2Department of Radiation Oncology, Seoul National University Hospital, Seoul, Korea, Republic of (South), 3Department of Radiology, Massachusetts General Hospital and Harvard Medical School, Boston, MA

Purpose/Objective(s):

We hypothesized that a verifiable LLM agent–orchestrated pipeline can automate radiotherapy treatment planning data extraction by linking OIS text-to-SQL cohort queries with TPS ESAPI-based DVH extraction while producing outputs that exactly match human-authored reference queries and scripts.

Materials/Methods:

The LLM agent pipeline was implemented using Apriel-15B-Thinker via Ollama on an NVIDIA A6000 GPU with 48 GB of memory and evaluated in the ARIA and Eclipse v18.0 environments. Incoming natural-language prompts were automatically classified into three query types: Type A (clinical intent SQL for cohort identification and treatment metadata by date range, disease site, technique, and treatment unit), Type B (ESAPI-based DVH metric extraction into tabular dataset), and Type C (Type A ? B: SQL-defined cohort followed by DVH extraction).

Type A was tested using 100 questions spanning 10 linguistic categories. Type B using 50 DVH scenarios involving PTVs and multiple OARs. Type C was evaluated in 10 end-to-end test cases that identified lung SABR patients treated within a predefined period and compiled DVH-derived metrics into a tabular dataset.

Example prompts included: Type A “Please provide a list of patients treated with lung SABR on a TrueBeam during 2025.”, Type B “For patient IDs A and B, please extract and tabulate DVH metrics for the ipsilateral lungs: V20Gy, D1000cc and mean dose.”, Type C “For patients treated with lung SABR on a TrueBeam in May 2025, please identify the cohort and then extract and tabulate ipsilateral lung DVH metrics: V20Gy, D1000cc, and mean dose.””

Outputs were normalized to a consistent schema to enable direct comparison across runs and scenarios. Human-authored SQL queries and ESAPI scripts generated reference outputs. A pass was defined as exact result-set agreement with the reference outputs (identical cohort membership/row set and identical returned fields/values); pass rates were computed per type (A–C).

Results:

All prompts were successfully processed. Type A achieved a 99% pass rate with consistent performance across linguistic categories, including negation/exclusion and relative time expressions. Type B achieved a 100% pass rate and reproduced all reference DVH metrics. Type C achieved a 90% pass rate; the single failure occurred because the requested OAR lacked a corresponding contoured structure name in the plan. Reliability was supported by structured parameter extraction, a two-step query planning/construction workflow, rule-based validation, and LLM-based semantic auditing with automated retry, which together detected and corrected errors before execution.

Conclusion: This verifiable LLM agent–orchestrated pipeline can reliably integrate OIS and TPS planning data into analysis-ready tables, reducing manual effort for cohort building and DVH harvesting to support outcomes research and quality improvement. Broader variable coverage is feasible through schema grounding on the full OIS SQL schema.