3142 - Redefining the Gold Standard: Benchmarking Commercial AI Models against Inter-Observer Variability in Complex Head and Neck OAR Contouring
Presenter(s)
C. T. Su1,2, S. A. Yeh2,3, S. Y. Hsu4, L. Y. Chang5, and Y. S. Liang6; 1Department of Radiation Oncology,E-Da hostpital, Kaohsiung City, Taiwan, 2Department of Medical Imaging and Radiological Sciences, I-Shou University, Kaohsiung City, Taiwan, 3Department of Radiation Oncology, E-Da hospital, Kaohsiung City, Taiwan, 4Department of Information Engineering, I-Shou University, Kaohsiung City, Taiwan, 5Department of Medical Imaging and Radiological Sciences, I-Shou University, Kaohsiung, Taiwan, 6Department of Radiation Oncology, E-Da hostpital, Kaohsiung, Taiwan
Purpose/Objective(s): To systematically evaluate the performance of two commercial deep-learning auto-contouring (DLAC) systems against an expert gold standard in head and neck (H&N) radiotherapy. While volumetric metrics are standard, they often fail to capture the clinical effort required for correction. This study integrates surface-based metrics to determine if AI-generated contours fall within the clinically acceptable bounds of observer variability, effectively redefining the efficiency benchmark for complex organs at risk (OARs).
Materials/Methods: Thirty H&N cancer patients treated with radiotherapy were retrospectively analyzed. Ground truth contours for over 10 OARs were manually delineated by a senior radiation oncologist with >20 years of experience, adhering to RTOG guidelines. Two commercial DLAC systems, EFAI AutoSeg (Vendor E) and AccuContour v4.0 (Vendor M), were used to generate auto-contours. Evaluation utilized standard geometric metrics (Volumetric Dice Similarity Coefficient [vDSC], 95% Hausdorff Distance [HD]) and surface-based efficiency metrics (Surface DSC [sDSC], Added Path Length [APL]). APL was specifically prioritized as a quantitative proxy for the "human-in-the-loop" correction burden, distinguishing between clinically negligible deviations (akin to inter-observer noise) and significant errors requiring substantial manual intervention.
Results: Both systems demonstrated high geometric concordance for large, high-contrast structures, achieving median vDSC >0.8 and HD <5mm, suggesting AI performance in these regions is statistically comparable to human expert consistency. Regarding the human-in-the-loop efficiency, the APL analysis revealed that for the majority of structures, the correction burden was minimal. Although variability was observed in select low-contrast regions, the APL metric provided a critical differentiation: it quantified that the actual manual refinement required was often localized, rather than necessitating global re-contouring. Comparative analysis indicated that Vendor E achieved statistically superior sDSC and lower APL in complex structures such as the Brainstem and Spinal Cord. This suggests that using APL as a proxy for correction burden effectively identifies where AI models align with the "Gold Standard" of efficiency, minimizing the time-cost of manual review even when geometric metrics suggest minor deviations.
Conclusion: Commercial DLAC systems have redefined the efficiency baseline for H&N planning, effectively emulating expert-level delineation for bony and large soft-tissue structures. The integration of APL serves as a vital proxy for clinical workload, demonstrating that even in complex OARs where geometric variability exists, the actual human-in-the-loop correction burden can remain manageable. These findings highlight that evaluating AI utility requires shifting focus from pure geometric overlap to quantifying the practical efficiency of physician oversight.