3796 - Domain-Specific Self-Supervised Pre-Training on Large-Scale MRI Data Enhances Organs-at-Risk Segmentation for MR-Linac Adaptive Radiotherapy In Cervical Cancer
Presenter(s)
Y. Wang1, W. Chen2, and C. Wang3; 1Shandong First Medical University and Shandong Academy of Medical Sciences, Taian, Shandong, China, 2Shandong First Medical University and Shandong Academy of Medical Sciences, Tianan ,Shandong, China, 3Department of Gynecologic Oncology, Shandong Cancer Hospital and Institute, Shandong First Medical University and Shandong Academy of Medical Sciences, Jinan, China
Purpose/Objective(s):
Cervical cancer is the fourth most common malignancy in women worldwide. Image-guided adaptive radiotherapy (ART) is the standard treatment for locally advanced cervical cancer. Accurate delineation of organs at risk (OARs) is crucial for reducing toxicity and optimizing dose, but ART requires MRI scans and OAR re-delineation within minutes on the treatment couch. Manual delineation is time-consuming and variable, limiting clinical throughput. We developed a deep learning model using a two-stage paradigm: first, large-scale unlabeled T1-weighted MRI for MAE self-supervised pre-training to extract anatomical features and build a pelvic MRI domain-specific model; second, transfer learning with limited labeled data for precise subvoxel OAR segmentation, meeting online clinical requirements.Materials/Methods:
We retrospectively collected over 5,000 unlabeled cervical cancer/pelvic MRI images for MAE self-supervised pre-training to explore anatomical common features, and over 100 annotated images for fine-tuning to achieve precise segmentation for adaptive radiotherapy. We adopted an asymmetric encoder-decoder MAE architecture: the encoder uses ViT structure processing visible unmasked patches with positional encoding, while the decoder uses a lightweight Transformer to reconstruct masked patches. For downstream segmentation, we constructed a dedicated Encoder-Decoder, freezing the first two layers to retain low-level features and fine-tuning the remaining layers end-to-end. Using five-fold cross-validation, data augmentation, and deep supervision, we achieved precise lesion segmentation. Diagnostic indicators including Dice coefficient (DC), average symmetric surface distance (ASSD), volume overlap error (VOE), and Hausdorff 95 distance (HD95) were calculated to evaluate prediction accuracy.Results:
In the clinical validation, the model based on MAE pre-training achieved significantly higher segmentation accuracy and stability in cervical cancer target region segmentation (DC: 76.82%, ASSD: 1.84mm, VOE: 26.85%, HD95: 8.10mm) compared to nnU-Net (DC: 74.35%, ASSD: 2.15mm, VOE: 29.42%, HD95: 10.11mm), demonstrating the superiority of this mask autoencoder-based pre-training strategy.Conclusion:
Our research has verified the clinical value of MAE self-supervised pre-training in the ART of cervical cancer. Through pre-training with over 5,000 unlabeled data, the model learned the common features of pelvic anatomy, effectively alleviating the problem of scarce labeled data; the fine-tuning with over 100 examples of precise-labeled data achieved task-specific optimization. The two-stage strategy balanced the breadth of representation learning and the accuracy of the segmentation task, achieving the optimal balance between annotation efficiency and segmentation performance compared to full-supervised learning.