Main Session
Sep
28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology
2579 - Anatomical Rule-Based Pipeline for AI-Assisted Quality Assurance of Breast Radiotherapy Contours
Presenter(s)
Tanguy Perennec, MD - ICO Nantes, Nantes, Nantes
T. Perennec1, G. Le Quellenec1, L. Vaugier1, D. Autret2, L. M. Sauvage1, A. Cailleteau1, A. Camps Malea1, F. Tomaszewski1, A. Mervoyer1, M. Le Blanc-Onfroy1, M. Legrand2, P. Guiouillier2, Q. Josset2, A. Moignier1, and N. Wagneur2; 1Institut de Cancérologie de l'Ouest René Gauducheau, France, Saint Herblain, France, 2Institut de Cancérologie de l'Ouest Paul Papin, France, Angers, France
Purpose/Objective(s):
Deep learning autosegmentation models are trained on manual contours and typically optimized using geometric similarity losses rather than explicit clinical constraints. As a result, they may produce subtle but clinically meaningful errors that can be missed during routine review, including boundaries inconsistency and continuity defects. We developed an automated quality-control (QC) pipeline based on rule-based spatial relationships between target volumes and reference anatomy and evaluated its agreement with double radiation oncologist review.Materials/Methods:
Breast wall targets, nodal clinical target volumes (CTVs), esophagus, and spinal canal were autosegmented using Raystation autosegmentation module (v1) on 30 female thoracic planning CT scans. The QC pipeline applied structure-specific rules derived from ESTRO breast contouring guidelines, based on: (i) relationships between target contours (e.g., laterality consistency), and (ii) relationships between target contours and reference anatomy, automatically derived with TotalSegmentator v2.5.0 (e.g., internal mammary node (IMN) caudal extent relative to the 4th rib). For each structure–rule pair and each scan, the algorithm produced a binary outcome (pass/fail). Outputs were compared with an independent assessment by two blinded radiation oncologists, adjucated by a third one. Check-level performance was summarized using accuracy, precision (PPV), and recall (sensitivity).Results:
Across 1,763 check-level evaluations, the QC algorithm achieved an overall accuracy of 0.860, precision of 0.870, and recall of 0.962. Continuity checks showed high agreement (accuracy 0.817–0.950), including 0.90 for esophagus and spinal canal. Lateral border checks demonstrated consistently high concordance (accuracy 0.813–0.967), with recall of 1.0 for all laterality rules (no false negatives). Dorso-ventral rules also performed well (accuracy 0.81–0.93). Cranio-caudal rules showed more variable performance (accuracy 0.55–0.91), with lower agreement for cranial axillary level I (accuracy 0.71) and cranial IMN (accuracy 0.55). These discrepancies are consistent with (i) limitations of some landmarks/rules (e.g., guidance combining multiple landmarks such as vein and artery), (ii) pixel-level differences without clinical impact, and (iii) higher inter-observer variability for borderline cranio-caudal endpoints. Concordance between two independent radiation oncologists was 0.75 and 0.73 for cranial axillary level I and cranial IMN, respectively, versus 0.94 overall across checkpoints.Conclusion:
A rule-based anatomical QC pipeline reliably detected continuity, laterality and dorso-ventral borders errors in breast radiotherapy autosegmentation. The main limitation was cranio-caudal endpoint checking, likely driven by ambiguous anatomical endpoints and inter-observer variability.