2497 - Comparative Evaluation of Embedding Architectures for Predicting Pathologic Complete Response of Breast Cancer from DCE-MRI Pharmacokinetic Maps
Presenter(s)
M. Kanthamneni1, H. Liu2, K. Tsang3, N. Tatonetti3, and T. Dou1; 1Department of Radiation Oncology, Cedars-Sinai Medical Center, West Hollywood, CA, 2Department of Hematology and Cellular Therapy, Cedars-Sinai Medical Center, West Hollywood, CA, 3Department of Computational Biomedicine, Cedars-Sinai Medical Center, West Hollywood, CA
Purpose/Objective(s):
Accurate pathologic complete response (pCR) prediction during early treatment of neoadjuvant chemotherapy (NAC) in breast cancer is essential for tailoring surgical and systemic interventions. The DCE-MRI-derived pharmacokinetic maps capture critical tumor microvasculature and permeability information. To study the computing architectures that enable optimal feature extraction for prognostic modeling, we implemented three modern encoding algorithms: Convolutional Neural Network- (CNN), Transformer-, and hybrid of CNN and Transformer-based architectures for predicting pCR and evaluated their discriminative performances.Materials/Methods: A cohort of 111 patients who received NAC was curated from the publicly available Duke Breast Cancer Dataset. Quantitative pharmacokinetic maps (Ktrans, Ve, Kep) were generated from pre-treatment DCE-MRI image data using the Tofts model with a Parker population-based arterial input function (AIF) using publicly available 3D Slicer module. These pharmacokinetic maps along with the tumor volumes-of-interest (VOIs) were used as the input to the embedding architectures for predicting pCR. Tumor VOIs were localized based on patient-specific clinical annotations and resampled to an isotropic resolution of 1 mm³. To mitigate the class imbalance between pCR (n = 31) and non-pCR (n = 78), we implemented a balanced sampling strategies during training. We performed a comparative evaluation of three embedding architectures: ResNet-18 (CNN), Swin Transformer, and CvT (Hybrid) model. These models were trained using NVIDIA GB10 GPU on a DGX Spark System. Models were optimized using 5-fold cross-validation (80% training, 20% validation) and evaluated using Area Under the Receiver Operating Characteristic Curve (Mean AUC).
Results: Table 1 shows the comparison on model performance and computation times. Swin-Transformer architecture achieved the highest predictive performance with a Mean AUC of 0.799, requiring a total training time of 142 minutes. The 3D ResNet-18 achieved the second-highest performance with a Mean AUC of 0.674, while requiring the lowest total training time of 93 minutes. The CvT model yielded a Mean AUC of 0.654 with a training time of 110 minutes.
Conclusion: Transformer-based models can successfully predict pCR from pharmacokinetic maps. Specifically, the Swin-Transformer achieved a superior Mean AUC of 0.799. This approach demonstrates the superior performance of transformers in analyzing complex microvascular patterns and tumor heterogeneity for clinical decision support. Table 1: Comparison of Embedding Architectures for Pathologic Complete Response (pCR) Prediction
| Architecture | Type | Mean AUC | Parameters (in millions) | Total Training time (in minutes) |
| Swin Transformer | Transformer | 0.799 | 8.7 | 142 |
| 3D-Resnet | CNN | 0.674 | 33.3 | 93 |
| CvT 3D | Convolutional Vision Transformer (Hybrid) | 0.654 | 4.8 | 110 |