Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2517 - Clinical Validation of Loss-Based Training for Unified Multi-Organ Auto-Contouring from Heterogeneously Labeled CT Datasets

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 17
POSTER

Presenter(s)

Jongmin Lee, - Oncosoft, Seoul, SEO

J. Lee, Y. J. Kim, S. Lee, H. J. Chae, and J. S. Kim; Oncosoft Inc., Seoul, Korea, Republic of (South)

Purpose/Objective(s):

Automated organ segmentation for radiation therapy planning relies on deep learning, but available CT datasets are typically annotated for different subsets of organs at risk (OARs). Pseudo-labeling, where a pre-trained model generates surrogate labels for unannotated organs, is commonly used to unify such datasets but risks propagating errors into clinical contours. Loss masking restricts loss computation to annotated classes per sample, avoiding pseudo-labels entirely. Although established in computer vision, this technique has not been validated for clinical auto-contouring. This study evaluates loss-based training across four anatomical regions against individually trained models and pseudo-labeling.

Materials/Methods:

In loss-based training, each sample carries metadata specifying which classes are annotated, and the loss is computed only over those classes. Four merging experiments were conducted: (1) Thoraco-abdominal: three datasets (abdomen, chest, bowel; 489 train/74 val) covering 18 OARs including liver, kidneys, lungs, and bowel; (2) Prostate: two datasets covering prostate, seminal vesicles, penile bulb, and rectal spacer (4 classes); (3) Spine/OAR: two datasets covering esophagus, stomach, spinal cord, cauda equina, duodenum, and spinal canal (6 classes); (4) Pelvic: four datasets covering rectum, bladder, femoral heads, bowel, and lumbosacral plexus (12 classes). A 3D U-Net with Dice + cross-entropy loss was used. Oversampling addressed class imbalance. Mean Dice Similarity Coefficient (DSC) was compared against individually trained models (Inner) and pseudo-label training.

Results:

Mean DSC on shared classes across four experiments (Shared = existing, New = added via merging) (Table)

In the thoraco-abdominal experiment, pseudo-labeling achieved a mean DSC of 0.879, underperforming loss-based training in 13 of 18 OARs. Prostate merging preserved accuracy for prostate (0.904 vs 0.902) and seminal vesicles (0.830 vs 0.825) while adding penile bulb (0.820). Spine/OAR merging improved spinal cord (0.901 vs 0.884) and stomach (0.923 vs 0.905) while integrating spinal canal (0.934). Pelvic consolidation of 12 OARs from four datasets maintained comparable accuracy and reduced inference time by 14-17%.

Conclusion:

Loss-based training reliably consolidates heterogeneously labeled CT datasets into unified auto-contouring models across multiple anatomical regions. This architecture-independent approach matched or exceeded individually trained models without pseudo-labels, while enabling new OAR integration and reducing inference time. By relying exclusively on verified annotations, it is well-suited for clinical deployment where contour fidelity is essential.

Region

Shared / New

Inner

Loss-Based

Thoraco-abd.

18 / 0

0.890

0.890

Prostate

3 / 1

0.848

0.849

Spine / OAR

5 / 1

0.857

0.869

Pelvic

10 / 2

0.896

0.895