Main Session
Sep 27
SS 01 - Reading the Whole Patient: Faces, Genes, Images, and the Future of Predictive AI

102 - Multi-Institutional Development and Prospective Validation of a Vision-Language Model for Prostate MRI Segmentation and Prognostication

03:40pm - 03:50pm ET
Room 254

Presenter(s)

Martin King, MD, PhD Headshot
Martin King, MD, PhD - Brigham and Women's Hospital, Boston, MA

M. T. King1, D. D. Yang2, C. Breneman3, B. Thomsen4, M. Sayan5, Y. Wang6, S. Rastogi7, S. C. Kamran8, K. Salari9, Y. Yuan10, J. A. Efstathiou8, M. E. Taplin7, P. L. Nguyen11, A. V. DAmico4, and J. E. Leeman4; 1Department of Radiation Oncology, Brigham and Women’s Hospital, Dana-Farber Cancer Institute, Harvard Medical School, Boston, MA, 2Department of Radiation Oncology, Brigham and Women’s Hospital/Dana-Farber Cancer Institute, Boston, MA, 3Dana-Farber/Brigham and Women's Cancer Center, Boston, MA, United States, 4Department of Radiation Oncology, Mass General Brigham, Harvard Medical School, Boston, MA, 5Rutgers Cancer Institute of New Jersey, Department of Radiation Oncology, New Brunswick, NJ, 6Department of Radiation Oncology, Massachusetts General Hospital, Harvard Medical School, Boston, MA, 7Dana-Farber Cancer Institute, Boston, MA, 8Massachusetts General Hospital, Boston, MA, 9Department of Urology, Mass General Brigham, Harvard Medical School, Boston, MA, 10Department of Radiation Oncology, Columbia University Vagelos College of Physicians and Surgeons, New York, NY, 11Mass General Brigham Cancer Institute, Boston, MA

Purpose/Objective(s): Deep-learning vision (V) segmentation models for prostate mpMRI could be used for delineating microboost volumes for radiation therapy (RT) or providing prognostic information. However, segmented lesions may be discordant with radiology reports. We hypothesized that a vision-language (VL) model, which incorporated information from radiology reports, could improve segmentation and prognostic performance.

Materials/Methods:

We obtained prostate MRI images from patients who underwent primary RT or radical prostatectomy (RP) between 2010-2019 across two institutions (labeled 1 and 2). The dataset was divided into 80:20 training-test splits. The V model was based on a UNet. In the VL model, lesion information from the radiology report (e.g. location, PI-RADS score) was fused with encoder features from the V model utilizing contrastive learning. Training was performed with 5-fold cross-validation (CV). The ground truth was manual segmentations of PI-RADS 3-5 lesions. Segmentation performance was evaluated on held-out test sets utilizing the PI-CAI score ([average precision + AUROC)/2]). A prognostic model was created by feeding segmentation features from decoder stages along with clinical features (e.g.2025 NCCN stage, radiographic T-stage [rT-stage]) into a survival head, which predicted metastasis. This model was trained with CV across the same folds as the segmentation model in order to prevent data leakage. Prognostication performance on the combined held-out test set was evaluated utilizing the c-index and multivariable Cox analysis (MVA), adjusted for NCCN and rT-stage. Finally, the VL model was applied to separate cohort of patients who were enrolled on 4 prospective protocols of neoadjuvant androgen deprivation therapy prior to RP (NEO) at institution 1. Given fewer events, MVA was only adjusted for NCCN stage.

Results:

The CV dataset included 1386 and 1135 patients from institutions 1 and 2, respectively. Regarding segmentation performance on the held-out test sets, the VL model had better PICAI scores than the V model for institutions 1 (0.837 (standard deviation (SD) 0.004 across model folds) versus 0.634 (0.018); p = 0.004; N = 347) and 2 (0.857 (0.010) versus 0.658 (0.010); p = 0.004; N = 298). For the NEO test set (N = 103), the PICAI score for VL was 0.917 (0.017). Regarding prognostication for the combined held-out set (N = 645; 40 events; median follow-up 5.9 years), the c-index for the model was 0.830 [0.761; 0.899] versus 0.746 [0.662; 0.831] (p = 0.006) for NCCN stage. The score was associated with metastasis (adjusted hazard ratio (AHR) 1.56; 95% confidence interval 1.21, 2.00; p < 0.001) on MVA. For the NEO set (N = 103; 24 events; median follow-up 6.4 years), the c-index of 0.786 [0.701; 0.871] was greater than that for NCCN stage (0.567 [0.460; 0.673] (p < 0.001)). The score was associated with metastasis (AHR 1.64; 95% CI: 1.30, 2.06; p < 0.001) on MVA.

Conclusion: In this multi-institutional analysis, a VL model provided improved tumor segmentation and prognostication.