1080 - Validation of an Automated Segmentation and Planning Reference Generation System for Gynecologic HDR Brachytherapy
Presenter(s)
R. Marchant1, Y. Gonzalez2, Y. Ding3, X. Feng4, X. Jia5, K. V. Albuquerque2, and C. R. Nwachukwu1; 1Department of Radiation Oncology, University of California San Diego, La Jolla, CA, 2Department of Radiation Oncology, University of Texas Southwestern Medical Center, Dallas, TX, 3Carina Medical LLC, Ashburn, VA, 4Carina Medical LLC, Lexington, KY, 5Department of Radiation Oncology and Molecular Radiation Sciences, Johns Hopkins University, Baltimore, MD
Purpose/Objective(s):
Materials/Methods: AutoBrachy incorporates deep learningbased multi-class segmentation, automated applicator reconstruction with dwell generation, and linear penalty model-based inverse planning using TG-43 dose calculation. Segmentation performance was assessed using Dice similarity coefficient (DSC) and 95th percentile Hausdorff distance (HD95). CTV D90 and OAR D2cc were compared between algorithm and clinical plans using one-sided t-tests. Statistical significance was defined as p < 0.05.
Results: A total of 189 studies were used for training and cross-validation, and 202 studies (54 patients) were included for external validation across tandem-and-ovoid (T&O) and tandem-and-ring (T&R) cases. In T&O cases, mean DSC ranged 0.690.78 (CTV), 0.750.85 (bladder), and 0.580.74 (rectum), with lower agreement for sigmoid (0.360.59) and small bowel (0.260.61). HD95 was lowest for CTV (59 mm) and bladder (612 mm) and higher for bowel (2247 mm). Similar trends were observed in T&R cases. CTV D90 was maintained between clinical and algorithm plans. In T&O cases, algorithm plans demonstrated significantly lower bladder, rectum, sigmoid, and small bowel D2cc compared with clinical plans (p<0.05). In T&R cases, only small bowel D2cc was significantly lower (p<0.05) (Table 1). Failed cases (n=12 T&O; n=79 T&R) were due to segmentation errors preventing applicator digitization and optimization.
Conclusion: AutoBrachy demonstrated reproducible multi-institutional segmentation accuracy and consistent plan quality across applicator types. Algorithm plans achieved OAR D2cc values comparable to or lower than clinical plans, particularly in T&O cases. Lower bowel agreement and T&R configurations highlight areas for further refinement.
Table 1. Comparison of OAR D2cc values between clinical (ground truth) and algorithm-generated contours/plans across institutions. Plans were considered failed when segmentation inaccuracies prevented applicator digitization and inverse planning. *p<0.05; **p<0.01.
| Applicator | T&O | T&R | |||
| Contour Source | Metrics | ROIs | Institution A n=138 | Institution B n=45 | Institution B n=145 |
| Failed cases | 12 | 0 | 79 | ||
| Ground Truth | D90 | CTV | 6.508 | 7.359 | 7.609 |
| D2cc | Bladder | 3.937 | 5.134 | 4.563 | |
| Rectum | 2.912 | 3.408 | 1.974 | ||
| Sigmoid | 3.371 | 4.156 | 3.991 | ||
| Small Bowel | 2.823 | 3.642 | 3.303 | ||
| Algorithm | D90 | CTV | 6.508 | 7.359 | 7.609 |
| D2cc | Bladder | 3.718* | 4.512** | 4.992 | |
| Rectum | 2.371** | 2.160** | 2.257 | ||
| Sigmoid | 2.909** | 3.064** | 4.129 | ||
| Small Bowel | 2.372** | 2.230** | 2.666** | ||