Main Session
Sep 29
PQA 05 - Physics

3166 - Automating Proton PBS Treatment Planning for Head and Neck Cancers Using Policy Gradient-Based Deep Reinforcement Learning

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 31
POSTER

Presenter(s)

Qingqing Wang, PhD - University of California San Diego, La Jolla, CA

Q. Wang1, and C. Chang2; 1University of California San Diego, San Diego, CA, 2California Protons Cancer Therapy Center, San Diego, CA

Purpose/Objective(s): Proton pencil beam scanning (PBS) treatment planning for head and neck (H&N) cancers is a complex, time-consuming process that requires balancing numerous conflicting planning objectives. While deep reinforcement learning has been applied to automate planning for simpler sites, existing models face significant limitations, including poor scalability and limited flexibility. Furthermore, these models often rely on weighted linear combinations of clinical metrics for rewards, which fail to generalize to the multi-target, multi-prescription complexities of H&N cancers. This study proposes an automated treatment planning model utilizing the proximal policy optimization (PPO) algorithm within a policy gradient DRL framework. The goal is to develop a system capable of managing high-dimensional, continuous action spaces and utilizing a dose distribution-based reward function to achieve human-level planning performance for complex H&N cases.

Materials/Methods: The planning process is formulated as an optimization problem in which empirical rules generate multiple objectives for target volumes and organs-at-risk (OARs), and an in-house optimization engine (L-BFGS) computes spot monitor unit (MU) values based on these conflicting objectives. A Transformer-based actor-critic agent, trained via PPO, iteratively adjusts up to 76 objective parameters (weights and dose limits) in a continuous action space in a human-like manner. The agent learns adjustment policy under the guidance of a novel dose distribution-based reward function, in which adaptive scoring is applied to individual planning structures. To manage the "curse of dimensionality," a feature selection module dynamically identifies eight OARs per iteration for adjustment. The model was trained and validated using a dataset of 34 H&N patients and further tested on 26 liver patients to assess generalizability.

Results: Plans generated by the PPO-based model for H&N patients demonstrated improved OAR sparing while maintaining equal or superior target coverage compared to human-generated clinical plans. The system successfully managed up to four prescription levels and 50 planning structures simultaneously. Additionally, the model showed strong generalizability, successfully producing high-quality plans for liver cancer patients.

Conclusion: The proposed DRL model is the first policy gradient auto-planning system to achieve human-level performance in PBS planning for H&N cancers. By utilizing policy gradients and a continuous action space, the framework overcomes the scalability and flexibility constraints of previous Q-learning approaches. This automated approach not only produces plans comparable or superior to those of experienced human planners but also demonstrates the potential for cross-site generalizability in radiation therapy.