Main Session
Sep 29
PQA 05 - Physics

3157 - Assessing the Generalizability and Robustness of Deep-Learning Dose Prediction in Head-and-Neck Radiotherapy to Clinically Realistic CT Perturbations

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 34

Presenter(s)

Amit Sawant, PhD Headshot
Amit Sawant, PhD - University of Maryland, Baltimore, Maryland

N. Tripathi1, R. Chowdhury2, L. Ren3, A. Sawant3, and B. D. Vaishnav3; 1Department of Computer Science, New York University, New York, NY, 2University of Maryland Medical System, St. Joseph Medical Center, Baltimore, MD, 3Department of Radiation Oncology, University of Maryland, School of Medicine, Baltimore, MD

Purpose/Objective(s):

The primary goal of this work is to establish a framework for assessing the generalizability of a deep-learning dose-prediction model for head-and-neck patients by evaluating its robustness to clinically realistic CT perturbations. Because multi-institution deployment must accommodate differences in scanners, reconstruction settings, and imaging practices, we examine how variations in noise, HU calibration, low-frequency bias, spatial resolution, and metal-induced artifacts influence dose-prediction accuracy.

Materials/Methods:

A 3D U-Net with squeeze-and-excitation blocks was trained on the OpenKBP dataset (200 training, 100 validation, 40 test cases) using masked MAE loss with 4× PTV weighting. Robustness was assessed on 40 test patients across 26 CT conditions: one baseline and several severity levels for heteroscedastic noise (P1), HU shifts (P2), low-frequency bias (P3), anisotropic blur (P4), and dental streak artifacts (P5). Perturbation ranges were chosen to meet or exceed ACR CT-simulation QA thresholds. Metrics included dose-MAE and DVH-based errors vs. ground truth. Robustness was quantified using ?DVH and ?MAE, the absolute change from each patient’s baseline.

Results:

Intensity-based perturbations (noise, bias, artifacts) produced minimal changes in ?DVH and ?MAE. HU-accuracy shifts were stable through clinically plausible ranges, with major degradation only at extreme offsets. Spatial-resolution degradation produced the largest and most systematic increases in both metrics, reflecting sensitivity to loss of anatomical edge definition.

Conclusion:

The model showed strong robustness to intensity-based CT variability, with meaningful degradation only under extreme HU shifts. Spatial-resolution loss consistently produced substantial errors, indicating CT resolution is more critical for cross-institution generalization than intensity-based variation.

Summary Table of Perturbation Effects

GenAI Disclosure

ChatGPT, OpenAI Codex, and Google Gemini were used to assist with text editing, condensation, and code debugging. All scientific content was fully reviewed, validated, and approved by the authors.

ParameterACR Thr.PerturbationSeverity?DVH?MAE
Noise~12HUP1(s_soft,s_bone)8–160HU0.0–1.4%0.0–0.3%
Uniformity5–7HUP3(A)10–200HU0.0–0.4%0.0–0.2%
CT# Accuracy0±4HUP2(?_soft,?_bone)5–1000HU-0.6–11.2%0.0–5.0%
ResolutionspecP4(s_xy,s_z)0.25–4.0blur1.5–18.2%0.2–10.5%
Artifactno-biasP5(amp,streak#)150–1200HU0.3–0.9%0.0–0.0%