1108 - Evaluating Radiation Therapy Quality Assurance Metrics as Predictors of Clinical Endpoints in Head and Neck Cancer: A Multi-Outcome Machine Learning Analysis of RTOG 0522
Presenter(s)
D. Wang1, X. Tie2, S. H. Lee1, H. Geng3, R. Caruana4, H. Zhong5, K. Men6, D. I. Rosenthal7, J. J. Caudell8, A. K. Singh9, C. U. Jones10, S. Rudra11, S. Brule12, T. J. Galloway13, G. Shenouda14, P. R. Anne15, C. Cardenas16, Q. T. Le17, and Y. Xiao1; 1Department of Radiation Oncology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, 2University of Pennsylvania, Philadelphia, PA, 3University of Pennsylvania, Perlman School of Medicine, Philadelphia, PA, 4Microsoft Research, Redmond, WA, 5Department of Radiation Oncology, University of Pennsylvania, Philadelphia, PA, 6Department of Radiation Oncology, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China, 7Department of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX, 8Moffitt Cancer Center, Tampa, FL, 9Department of Radiation Medicine, Roswell Park Comprehensive Cancer Center, Buffalo, NY, 10Sutter Medical Center Sacramento, Roseville, CA, 11Department of Radiation Oncology, Winship Cancer Institute of Emory University, Atlanta, GA, 12The Ottawa Hospital Cancer Program, Ottawa, ON, Canada, 13Fox Chase Cancer Center, Philadelphia, PA, 14McGill University Health Center, Montreal, QC, Canada, 15Thomas Jefferson University, Philadelphia, PA, 16University of Alabama at Birmingham, Birmingham, AL, 17Stanford University, Stanford, CA
Purpose/Objective(s): To investigate the impact of radiation therapy quality assurance (RTQA) on multiple clinical outcomes in head and neck carcinoma using machine learning approaches.
Materials/Methods: We performed a retrospective analysis of 669 patients from RTOG 0522, a phase III trial in stage III-IV head and neck squamous cell carcinoma. Collected features included clinical variables (e.g., age, sex, tumor site, tumor size and stage, smoking history, and hemoglobin level), RTQA metrics including contour and dose-volume analysis (DVA) scores for target volumes and OARs, and DVH parameters for the PTVs and four protocol-specified OARs dose constraints. Patients were stratified into training (n=502) and independent testing (n=167) cohorts. Four outcomes were evaluated: local-regional recurrence (LR), disease-free survival (DFS), 2-year overall survival (OS), and treatment-related adverse events. We compared single-task XGBoost with Lasso feature selection against multi-task GBDT (MT-GBDT) using the union of Lasso-selected features. Model performance was evaluated using AUC with bootstrap 95% confidence intervals. The predictions of all models were interpreted using partial dependence analysis.
Results: Single-task models achieved AUCs of 0.654 (95% CI: 0.550-0.754) for LR using 17 features, 0.678 (0.593-0.751) for DFS using 22 features, 0.843 (0.773-0.910) for 2-year OS using 14 features, and 0.511 (0.503-0.664) for adverse events using 8 features. MT-GBDT using 24 union features showed numerically lower performance with overlapping confidence intervals, yielding with AUCs of 0.596 (0.479-0.712), 0.660 (0.576-0.738), 0.780 (0.668-0.879), and 0.537 (0.503-0.675) for the respective outcomes. Notably, RTQA metrics—specifically target volume contour scores and DVA scores—were consistently selected by Lasso across all four outcomes, though their relative importance varied by outcome. Other commonly selected features included cT/cN stage, age, hemoglobin level, parotid gland D50, and larynx mean dose. Partial dependence analysis confirmed that better contour and treatment plan quality correlated with improved outcomes, while higher clinical stages, increased OAR doses, older age, and lower hemoglobin were associated with unfavorable endpoints.
Conclusion:
This multi-outcome analysis demonstrates that machine learning models integrating RTQA metrics with clinical and dosimetric features are prognostic of diverse clinical endpoints in head and neck cancer. Single-task models outperformed multi-task learning, indicating feature-outcome relationships differ substantially across endpoints, necessitating individualized modeling approaches for the evaluated dataset. Due to the limited OAR-related data, prediction of adverse events remains challenging, highlighting the need for more comprehensive OAR contour and dosimetric information—particularly for toxicity-relevant structures—to enable more robust analyses.