3088 - A Multi-Platform Evaluation of Performance and Inter-Algorithm Variability for Organs at Risk: JASTRO 2024 Nationwide Auto-Segmentation Trial
Presenter(s)
H. Nemoto1,2, M. Saito1, S. Kito3, M. Matsuda1, T. Komiyama1, D. Kawahara4, R. Tozuka1,2, and H. Onishi1; 1Department of Therapeutic Radiology, University of Yamanashi, Chuo, Japan, 2Department of Radiation Oncology, Tohoku University Graduate School of Medicine, Sendai, Japan, 3Tokyo Metropolitan Cancer and Infectious Diseases Center Komagome Hospital, Tokyo, Japan, 4Department of Radiation Oncology, Hiroshima University Hospital, Hiroshima, Japan
Purpose/Objective(s):
The rapid clinical integration of auto-segmentation (AS) for organs at risk (OARs) promises significant workflow efficiency. This study presents the findings of a nationwide trial organized in 2024 by the 37th annual meeting of Japanese Society for Radiation Oncology. The purpose of this study was to evaluate the performance and inter-algorithm variability of current AS platforms to establish clinical quality assurance and safety standards.Materials/Methods:
Participation was open to all available AS platforms at each institution, including commercial deep-learning (DL), model-based, atlas-based, and in-house developed algorithms. Manual contour modification was strictly prohibited to evaluate the intrinsic performance of the algorithms. Two CT datasets as test data (Thorax and Head-and-Neck [HN]) were distributed to participants. Target OARs consisted of clinically common structures: 5 in the Thorax and 7 in the HN region. Ground truth was established by radiation oncologists primarily following RTOG/NRG Oncology guidelines. Performance and efficiency were quantified using the Dice Similarity Coefficient (DSC), 95th percentile Hausdorff distance (HD95), and processing time.Results:
A total of 24 unique platforms for the Thorax and 23 for the HN region were evaluated. These comprised 15 commercial DL platforms for both regions, alongside 9 and 8 other platforms (model-based, atlas-based, or in-house DL) for the Thorax and the HN region, respectively. While high-contrast thoracic OARs demonstrated high stability, represented by the Lungs (DSC range 0.94–0.98), substantial inter-platform discordance was observed in low-contrast or small-volume structures, particularly the Esophagus (0.46–0.91) and Optic Chiasm (0.10–0.47). Table 1 summarizes detailed performance and efficiency metrics for all platforms. Comparison across platforms revealed that commercial DL solutions consistently dominated the top rankings. Processing times varied across regions and platforms, showing no correlation with geometric accuracy.Conclusion:
These results show that although top tier AS platforms achieve high accuracy, inter-platform variability persists in clinical practice. These findings highlight the need for structure-specific validation and cautious clinical integration of automated contours. Table 1. Performance and Efficiency: Overall Trial Median vs. Top 3 Ranked Platforms| Region & Metric | Overall (Median [Range]) | Top 1 | Top 2 | Top 3 |
| Thorax (n=24) | (Platform A) | (Platform B) | (Platform C) | |
| Mean DSC | 0.88 [0.46 - 0.98] | 0.94 | 0.93 | 0.91 |
| Mean HD95 (mm) | 3.82 [1.62 – 57.44] | 2.41 | 2.87 | 2.94 |
| Mean Time (sec) | 101.50 [7.00 – 625.00] | 44 | 231 | 35 |
| H&N (n=23) | (Platform A) | (Platform D) | (Platform E) | |
| Mean DSC | 0.69 [0.10 - 0.93] | 0.77 | 0.73 | 0.70 |
| Mean HD95 (mm) | 4.14 [2.08 – 13.50] | 3.27 | 3.16 | 3.52 |
| Mean Time (sec) | 40.00 [12.00 – 242.00] | 37 | 17 | 25 |