1152 - Community-Based Health System Validation of the SHIELD-RT Machine Learning Model to Prevent Acute Care Events During Radiotherapy
Presenter(s)
M. V. Elia1, R. Benson2, N. Bhargava3, J. Levey4, N. Eclov5, I. Friesner6, S. C. D. Hampson6, P. H. Fuerst7, J. Alexander8, A. Witzum9, D. Spiegel10, M. Palta5, J. Feng11, E. J. Yoshida12, and J. C. Hong9; 1UCSF + UC Berkeley Joint Program in Computational Precision Health, San Francisco, CA, San Francisco, CA, 2Department of Radiation Oncology, University of California San Francisco, San Francisco, CA, 3Department of Radiation Oncology, Beth Israel Deaconess Medical Center, Boston, MA, 4Brigham and Women's Hospital/Dana Farber Cancer Institute, Boston, MA, United States, 5Duke University Medical Center, Department of Radiation Oncology, Durham, NC, 6University of California, San Francisco, Bakar Computational Health Sciences Institute, San Francisco, CA, 7Washington Hospital Radiation Oncology Center, Fremont, CA, 8University of California, San Francisco, San Francisco, CA, 9University of California, San Francisco, Department of Radiation Oncology, San Francisco, CA, 10Department of Radiation Oncology Brigham and Women's Hospital, Boston, MA, 11UCSF, San Francisco, CA, 12Department of Radiation Oncology, University of California San Francisco,, San Francisco, CA
Purpose/Objective(s):
Roughly 10-20% of patients undergoing radiation therapy (RT) require acute care via emergency visits or hospitalization. We previously reported the results of the System for High-Intensity Evaluation During Radiotherapy (SHIELD-RT), one of the first machine learning (ML)-guided randomized controlled trials in healthcare, where ML was applied to electronic health record (EHR) data to identify patients at high risk for acute care events and direct increased clinical evaluations – reducing acute care events by 45% and overall costs by 48%. We previously validated the SHIELD-RT model in two external academic health centers (AHCs), where performance was largely robust. Community health systems (CHSs) often serve a different patient population with different needs. Clinical ML models are typically developed and evaluated in AHCs, which can consequently impact performance in CHSs, where validation is uncommon. Thus, we sought to test the hypothesis that the SHIELD-RT model would maintain its performance in a CHS with distinct patient population, clinical practice, and EHR.Materials/Methods:
This IRB-approved study evaluated the SHIELD-RT model on 1,047 RT courses at a CHS (Jan 2013 - Mar 2022). For each course, the previously tested gradient boosted tree model was applied to structured EHR data including patient characteristics, cancer treatment, vitals, lab results, medications, and prior acute care utilization. Model performance was evaluated using AUROC, Brier score, and calibration plots. As in SHIELD-RT, we classified high-risk RT courses as those with an acute care risk prediction of >10%.Results:
The model demonstrated good performance with an AUROC of 0.755 (95% CI: [0.709, 0.801]), comparable to previous AHC results, with a performance drop from 0.851 in the non-intervention courses during SHIELD-RT. The sensitivity and specificity at the 10% threshold were 58.0% and 80.0%, (SHIELD-RT: 69.8% and 77.5%). Brier score was 0.08, corroborated by calibration plots demonstrating strong calibration, particularly at probabilities <40%. The model classified 23% of courses as high-risk in the CHS. The true event rates of the high- and low-risk populations were 23.3% and 5.0%, demonstrating that the discriminatory power of the model remains strong in a CHS.Conclusion:
We externally validated a clinically tested model to predict risk of acute care utilization among patients undergoing outpatient radiotherapy in a CHS. The model performance remained largely consistent with results from two external AHCs. As the SHIELD-RT study previously demonstrated the model’s ability to direct care and reduce acute care event rates, these results show promise for its generalizable clinical impact in a range of clinical settings. We are currently evaluating the model in an additional CHS and assessing sociodemographic subgroup performance across institutions.