Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2620 - Large Language Model with Spatial Mask Guidance and Automatic Gross Tumor Volume Delineation In Esophageal Cancer: A Multicenter Validation Study

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 20
POSTER

Presenter(s)

Lecheng Jia, PhD - Tianjin University, Tianjin, Tianjin

H. Sun1, Z. An2, L. Jia3, and L. N. Zhao4; 1Department of Radiation Oncology, Xijing Hospital, Air Force Medical University, xi'an, Shaanxi, China, 2Department of Radiation Oncology, Xijing Hospital, Fourth Military Medical University, Xi'an, Shaanxi, China, 3Shanghai United Imaging Healthcare Co., Ltd., Shanghai, China, 4Department of radiation oncology, Xijing Hospital, the Fourth Military Medical University, Xi'an, China

Purpose/Objective(s):

Accurate gross tumor volume (GTV) delineation is a cornerstone of radiotherapy planning for esophageal cancer. However, under CT-only conditions, indistinct tumor boundaries often lead to limited robustness of existing automatic delineation methods. This study aimed to develop a dual-prompt guided framework for automatic GTV delineation in esophageal cancer by integrating large language model (LLM) with spatial mask guidance.

Materials/Methods:

A retrospective multicenter cohort of 857 patients with esophageal cancer was included, comprising one internal cohort and two external cohorts. The proposed framework consists of two stages: (1) a CT-based automatic segmentation network generating an initial GTV as a spatial mask prompt, and (2) a refinement network that jointly incorporates spatial mask prompt and LLM-generated text prompt derived from structured clinical information. Performance was evaluated using overlap, surface distance, and volumetric metrics, and compared with multiple state-of-the-art deep learning models. Ablation experiments were conducted to assess the contribution of two prompts. Radiomics analyses were further performed to evaluate the clinical equivalence between automatically delineated and manual GTVs.

Results:

On the internal cohort, the proposed model achieved DSC of 80.32±6.65 %, significantly outperforming all comparison methods, with lower HD95 (11.42±8.90 mm) and absolute volume differences (9.99±10.49 ml). Consistent performance advantages were maintained across both external cohorts. Ablation results demonstrated that the combined use of text and spatial mask prompts yielded superior accuracy and boundary stability compared with text-only or no-prompt settings. Radiomics analyses showed high feature robustness and strong correlations between radscores derived from automatic and manual GTVs, with no significant differences in prognostic performance.

Conclusion:

The proposed dual-prompt guided framework improves the accuracy and cross-center generalizability of CT-only automatic GTV delineation for esophageal cancer, providing a reliable foundation for radiotherapy planning.