3564 - Beyond Overlap Metrics: Uncertainty-Based Triage for Pathology-Anchored QA of Nodal Autosegmentation In HPV-Positive Oropharyngeal Cancer
Presenter(s)
Y. I. Mohamed1, M. M. Badawy1, A. Ali1, C. Sjogreen1, A. S. Mohamed1, K. Elsayes2, M. T. Spiotto1,3, G. B. Gunn1, A. Lee1, S. Y. Lai1, A. C. Moreno1, K. A. Hutcheson4, S. J. Frank1, A. S. Garden1, D. I. Rosenthal1, Z. Ye5, R. Mojahed-Yazdi6, B. H. Kann7, C. D. Fuller8, and M. Naser1,8; 1Department of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX, 2MD Anderson Cancer Center, Houston, TX, 3Department of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX, USA, houston, TX, 4The Department of Head and Neck Surgery, The University of Texas MD Anderson Cancer Center, Houston, TX, 5Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School, Boston, MA, 6Brigham and Women's Hospital, Fall River, MA, United States, 7Department of Radiation Oncology, Mass General Brigham/Dana-Farber Cancer Institute, Harvard Medical School, Boston, MA, 8Division of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX
Purpose/Objective(s):
Missed or spurious nodal targets can risk geographic miss or unnecessary dose. We hypothesized that ensemble-derived uncertainty can flag low-quality nodal GTV (GTVn) autosegmentation for selective review, and that adding FDG-PET to contrast-enhanced CT (CECT) improves uncertainty-based failure triage.Materials/Methods:
We curated 131 HPV-positive oropharyngeal cancer cases with 187 pathology-confirmed metastatic lymph nodes (FDG-PET in 75/131). Three experts contoured each node; STAPLE consensus of pathology-confirmed nodes served as reference (median expert-to-consensus Dice similarity coefficient [DSC] 0.95). Three 3D nnU-Net ensembles were trained: DS001 CECT-only (131-case cohort), DS002 CECT-only (PET subset, n=75), and DS003 PET+CECT (PET subset, n=75). Held-out tests were n=21 (DS001) and n=10 (PET cohort; DS002/DS003). Uncertainty (coefficient of variation [CV], mean absolute deviation ratio [MADR], ROI ensemble entropy) within predicted GTVn was evaluated by Spearman correlation with DSC and area under the receiver operating characteristic curve (AUROC) for failures (DSC<0.90). A selective-review simulation assumed correction of the 20% most uncertain PET test cases (2/10) to expert contours.Results:
Mean DSC was 0.835 for DS001 (n=21) and 0.939 and 0.938 for DS002 and DS003 (n=10), respectively. Best-performing uncertainty metric differed by dataset: DS001 CV rho=-0.668 (p=0.0009); DS002 ROI entropy rho=-0.721 (p=0.019); DS003 ROI entropy rho=-0.939 (p=0.0001). AUROC for DSC<0.90 failures was 0.90 (DS001; MADR), 0.94 (DS002), and 1.00 (DS003), with wide confidence intervals due to small n. In the PET cohort, reviewing 2/10 cases raised the minimum unreviewed DSC from 0.853 to 0.884 (DS002) and from 0.868 to 0.928 (DS003); DS003 triaged both failures (DS002 triaged 1 of 2).Conclusion:
Ensemble uncertainty provides an actionable QA signal for targeted triage and selective review of small, pathology-confirmed nodal targets. In this pilot PET cohort, PET+CECT improved uncertainty-based failure triage while maintaining similar mean overlap, supporting workflow-compatible QA for deployment.