Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2497 - Comparative Evaluation of Embedding Architectures for Predicting Pathologic Complete Response of Breast Cancer from DCE-MRI Pharmacokinetic Maps

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 17
POSTER

Presenter(s)

Mohith Kanthamneni, MS - Department of Radiation Oncology, Cedars-Sinai Medical Center, Los Angeles, CA

M. Kanthamneni1, H. Liu2, K. Tsang3, N. Tatonetti3, and T. Dou1; 1Department of Radiation Oncology, Cedars-Sinai Medical Center, West Hollywood, CA, 2Department of Hematology and Cellular Therapy, Cedars-Sinai Medical Center, West Hollywood, CA, 3Department of Computational Biomedicine, Cedars-Sinai Medical Center, West Hollywood, CA

Purpose/Objective(s):

Accurate pathologic complete response (pCR) prediction during early treatment of neoadjuvant chemotherapy (NAC) in breast cancer is essential for tailoring surgical and systemic interventions. The DCE-MRI-derived pharmacokinetic maps capture critical tumor microvasculature and permeability information. To study the computing architectures that enable optimal feature extraction for prognostic modeling, we implemented three modern encoding algorithms: Convolutional Neural Network- (CNN), Transformer-, and hybrid of CNN and Transformer-based architectures for predicting pCR and evaluated their discriminative performances.

Materials/Methods: A cohort of 111 patients who received NAC was curated from the publicly available Duke Breast Cancer Dataset. Quantitative pharmacokinetic maps (Ktrans, Ve, Kep) were generated from pre-treatment DCE-MRI image data using the Tofts model with a Parker population-based arterial input function (AIF) using publicly available 3D Slicer module. These pharmacokinetic maps along with the tumor volumes-of-interest (VOIs) were used as the input to the embedding architectures for predicting pCR. Tumor VOIs were localized based on patient-specific clinical annotations and resampled to an isotropic resolution of 1 mm³. To mitigate the class imbalance between pCR (n = 31) and non-pCR (n = 78), we implemented a balanced sampling strategies during training. We performed a comparative evaluation of three embedding architectures: ResNet-18 (CNN), Swin Transformer, and CvT (Hybrid) model. These models were trained using NVIDIA GB10 GPU on a DGX Spark System. Models were optimized using 5-fold cross-validation (80% training, 20% validation) and evaluated using Area Under the Receiver Operating Characteristic Curve (Mean AUC).

Results: Table 1 shows the comparison on model performance and computation times. Swin-Transformer architecture achieved the highest predictive performance with a Mean AUC of 0.799, requiring a total training time of 142 minutes. The 3D ResNet-18 achieved the second-highest performance with a Mean AUC of 0.674, while requiring the lowest total training time of 93 minutes. The CvT model yielded a Mean AUC of 0.654 with a training time of 110 minutes.

Conclusion: Transformer-based models can successfully predict pCR from pharmacokinetic maps. Specifically, the Swin-Transformer achieved a superior Mean AUC of 0.799. This approach demonstrates the superior performance of transformers in analyzing complex microvascular patterns and tumor heterogeneity for clinical decision support. Table 1: Comparison of Embedding Architectures for Pathologic Complete Response (pCR) Prediction

Architecture Type Mean AUC Parameters (in millions) Total Training time (in minutes)
Swin Transformer Transformer 0.799 8.7 142
3D-Resnet CNN 0.674 33.3 93
CvT 3D Convolutional Vision Transformer (Hybrid) 0.654 4.8 110