Main Session
Sep 27
SS 01 - Reading the Whole Patient: Faces, Genes, Images, and the Future of Predictive AI

103 - Predicting Survival from Appearance: A Multimodal Deep Learning Model Integrating Facial Photographs and MRI

03:20pm - 03:30pm ET
Room 254

Presenter(s)

Zihang Chen, MD Headshot
Zihang Chen, MD - University of Texas Southwestern Medical Center at Dallas, Dallas, TX

Z. Chen1, S. Zhang1, M. Hao1, X. Han2, Z. Wang2, and Y. Sun1; 1Department of Radiation Oncology, Sun Yat-sen University Cancer Center, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Guangdong Key Laboratory of Nasopharyngeal Carcinoma Diagnosis and Therapy, Guangzhou, China, 2School of Biomedical Engineering, Shanghai Jiaotong University, Shanghai, China

Purpose/Objective(s):

Facial appearance reflects systemic health status, whereas MRI captures detailed local tumor characteristics. We hypothesized that integrating external phenotypic features derived from facial photographs with internal tumor information from MRI would improve survival prediction and risk stratification in patients with nasopharyngeal carcinoma. To test this, we developed a multimodal deep learning model combining facial photographs and MRI to predict survival risk.

Materials/Methods:

Pretreatment magnetic resonance imaging (MRI) scans and facial photographs obtained prior to radiotherapy were retrospectively collected for 1,250 patients with nasopharyngeal carcinoma from two institutions. Three deep learning survival prediction models were developed: a Face model, an MRI model, and a multimodal Face–MRI model. Models were trained in 821 patients, internally validated in 206 patients, and externally tested in an independent cohort of 223 patients.For facial photographs, DINOv3, a state-of-the-art vision foundation model, was used as an appearance encoder to extract survival-relevant surface phenotypic features (Face model). For MRI, a ResNet architecture served as an anatomy encoder to capture complex internal structural and pathological characteristics (MRI model). To integrate complementary information across modalities, a Mixture-of-Experts architecture was implemented to construct the multimodal Face–MRI model.

Results:

The MRI and Face models demonstrated comparable predictive performance across all datasets. The MRI model achieved C-indices of 0.858, 0.712, and 0.693 in the training, internal validation, and external validation cohorts, respectively, while the Face model achieved C-indices of 0.861, 0.724, and 0.711. The multimodal Face–MRI model further improved predictive performance, with C-indices of 0.879, 0.744, and 0.721 across the three cohorts. The Face–MRI model also showed robust performance in risk stratification.Kaplan–Meier analysis demonstrated superior 3-year overall survival in the low-risk group compared with the high-risk group. Among high-risk patients, those receiving immunotherapy in addition to standard treatment exhibited improved survival relative to those treated with standard therapy alone. In contrast, among low-risk patients, the addition of immunotherapy did not confer a survival benefit.

Conclusion:

A multimodal deep learning model integrating facial photographs and MRI provides accurate survival prediction and effective risk stratification in nasopharyngeal carcinoma. Both modalities are noninvasive and readily obtainable in routine clinical practice, with facial photographs offering complementary information reflecting systemic health status. This approach may enable noninvasive identification of patients most likely to benefit from immunotherapy and support personalized treatment decision-making.