368 - A Modular Multi-Agent Reinforcement Learning Framework Guided by LLMs: Improving Quality and Efficiency of Treatment Planning in Lung Radiotherapy
Presenter(s)
Z. Wang, H. Guo, Y. Lei, R. Samstein, K. Rosenzweig, M. Chao, T. Liu, J. Xia, and J. Zhang; Department of Radiation Oncology, Icahn School of Medicine at Mount Sinai, New York, NY
Purpose/Objective(s):
Reinforcement learning (RL) has shown promise in automated treatment planning due to its ability to model long-term action–reward relationships. However, training full RL agents specifically designed for individual planning tasks is unrealistic for widespread clinical adoption because of high training cost and poor generalizability. Large language models (LLMs), on the other hand, provide decision-making capability but lack explicit long-term optimization awareness and often require multiple iterative refinements without guaranteed optimality. In this work, we propose a hybrid LLM-guided RL framework that integrates the strengths of both approaches for robust and efficient treatment plan optimization.
Materials/Methods:
Instead of training a monolithic RL agent for the entire planning workflow, we decompose the problem into smaller, well-defined optimization sub-tasks, each handled by a task-specific RL agent. Dedicated Q-learning RL agents were trained using linear function approximation for solving parallel OAR, serial OAR, and PTV coverage optimizations. Six representative cases independent of the evaluation cohort were used to train the RL agents. An LLM-based supervisory agent operated at a higher level, interpreting dosimetric feedback and selectively activating appropriate RL agents. The LLM agent does not directly modify plan parameters but coordinates the sequence and scope of RL-driven actions. GPT-4o (OpenAI) was used as the underlying LLM. Sixty-two retrospective locally advanced NSCLC cases were evaluated using institutional knowledge-based planning (KBP) plans as optimization starting points. All plan refinements were performed in Eclipse via ESAPI. Clinical goal achievement rates were compared among the initial KBP plans, an LLM-only automated planning system, and the proposed hybrid system.Results:
The hybrid framework increased the proportion of plans meeting all clinical constraints from 73% (KBP) to 94% (LLM-RL), comparable to the LLM-only system. It yielded lower lung V20 than the LLM-only system (28.73 ± 10.1% vs. 29.62 ± 10.6%, p = 0.048), with a maximum per-case improvement of 4.31%. No significant differences were observed in other dosimetric metrics. The hybrid framework also improved optimization efficiency and reduced the mean number of refinement steps from 6.2 to 3.2 (48% reduction).
Conclusion:
A hybrid LLM–RL architecture provides a practical balance between long-term optimization capability and high-level decision flexibility. By restricting RL training to structured sub-tasks and positioning the LLM as a supervisory controller, the proposed framework reduces training complexity, improves OAR sparing, and nearly halves the number of refinement steps. This approach demonstrates a scalable pathway toward clinically deployable autonomous radiotherapy treatment planning systems.