2414 - Quality over Quantity: Impact of Expert-Refined Radiotherapy-Specific MRI on Deep Learning-Based Urethra Segmentation
Presenter(s)
Y. Bai1,2, Q. Chen3, Y. Rong1, and Z. Libing1; 1Department of Radiation Oncology, Mayo Clinic AZ, Phoenix, AZ, 2Department of Radiation Oncology, Peking University First Hospital, Beijing, China, 3Department of Radiation Oncology, Mayo Clinic, Phoenix, AZ
Purpose/Objective(s): Urethra-sparing is vital for reducing genitourinary toxicity in prostate radiotherapy (RT). However, the urethra is frequently obscured on CT and inconsistent on diagnostic MRI. Clinical RT workflows introduce further complexities, such as androgen deprivation therapy (ADT) induced atrophy, perirectal spacers, and rectal balloons, which significantly degrade the performance of models trained on diagnostic-only data. This study evaluates the efficacy of the Carina AI platform trained on multi-institutional, RT-specific refined datasets compared to large-scale diagnostic repositories.
Materials/Methods: Three training cohorts were assessed: Group A (n=200): Diagnostic MRI from the public TCIA-prostate-zones database with original labels. Group B (n=40): High-quality refined data consisting of 20 expert-corrected TCIA cases and 20 institutional RT-MRI cases (Mayo Clinic). Group C (n=240): Combined dataset (A+B). The models were evaluated on an independent test set of 20 Mayo Clinic RT-MRI cases. To enhance structural robustness, post-processing via Largest Connected Component (LCC) analysis was implemented. Performance was quantified using Dice Similarity Coefficient (DSC), Percent Coverage, and 95th Percentile Hausdorff Distance (HD95).
Results: Group B demonstrated superior performance with a median DSC of 0.59 and HD95 of 5.8 mm, significantly outperforming both Group A (the diagnostic cohort) and historical benchmarks (median DSC ~0.40). Notably, Group C (n=240) yielded inferior results compared to Group B (n=40), suggesting that the inclusion of 200 unrefined diagnostic cases introduced substantial label noise and class imbalance, misleading the AI’s spatial convergence. While Group B achieved high anatomical precision (DSC), its percent coverage (60%) was lower than dilated historical models (81-89%), reflecting a prioritized focus on true anatomical boundaries rather than artificial safety margins.
Conclusion: High-quality, expert-refined datasets are superior to large-scale, unrefined diagnostic data for specialized OAR segmentation in radiotherapy. Our results indicate that the Carina AI platform can achieve clinically superior precision even with small-sample training, provided the data accounts for RT-specific variables like spacers and balloons. This paradigm shift towards "Data-Centric AI" is essential for reliable OAR sparing in complex clinical scenarios.