Main Session
Sep 29
PQA 07 - Head and Neck Cancer, Lung Cancer/Thoracic Malignancies, and Nursing and Supportive Care

3696 - A Two-Stage Latent Diffusion Model for Predicting FBCT Images During Radiotherapy: Enabling Adaptive Radiotherapy Trigger Monitoring with Reduced Scanning Frequency

03:45pm - 05:00pm ET
Poster Hall - Exhibit Hall A
Screen: 9
POSTER

Presenter(s)

Lecheng Jia, PhD - Tianjin University, Tianjin, Tianjin

Z. Liu1, M. Hao2, Z. Mo3, L. Jia4, and G. Q. Zhou2; 1Southern Medical University, Guangzhou, China, 2Department of Radiation Oncology, Sun Yat-sen University Cancer Center, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Guangdong Key Laboratory of Nasopharyngeal Carcinoma Diagnosis and Therapy, Guangzhou, China, 3Shenzhen United Imaging Research Institute of Innovative Medical Equipment, Shenzhen, China, 4Shanghai United Imaging Healthcare Co., Ltd., Shanghai, China

Purpose/Objective(s):

This study aims to develop a two-stage latent diffusion model for predicting future fan-beam computed tomography (FBCT) images during radiotherapy for nasopharyngeal carcinoma (NPC), and to propose an optimized, model-assisted FBCT scanning protocol, with the goal of enabling adaptive radiotherapy (ART) trigger monitoring.

Materials/Methods:

Retrospective data from 464 NPC patients (8,880 FBCT scans during radiotherapy) were randomly split into training (64%), validation (16%), and test (20%) sets. Clinical features included TNM stage, age, sex, and time relative to radiotherapy initiation.

FBCT Scanning Protocol: A treatment week comprised five radiotherapy sessions. For weeks 3, 4, and 5 (ART-indicated weeks), the model was designed to predict the last two FBCT sessions using the first three real scans as input.

A two-stage latent diffusion model was developed. First, a 3D autoencoder was pre-trained on the training set to extract latent space features from FBCT scans. Second, the diffusion model was trained sequentially: in the first phase, an upstream task synthesized FBCT images from clinical features; in the second phase, a downstream task integrated a multi-temporal FBCT-guided control network into the backbone to generate future FBCT images.

Performance evaluation included: (1) image quality assessed using Mean Absolute Error (MAE), Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR) and Normalized Cross-Correlation (NCC); (2) auto-segmentation of organs at risk (OARs) evaluated using Dice Similarity Coefficient (DSC); (3) dose validation performed by comparing dose distributions between predicted and real scans, with results expressed as absolute error.

Results:

For image quality assessment, the test set metrics were: MAE 9.0 ± 1.5 (HU), SSIM 0.9734 ± 0.0071, PSNR 34.79 ± 1.42, NCC 0.9906 ± 0.0035. Representative DSC values for OAR auto-segmentation on predicted images were: brainstem 0.943 ± 0.01, temporal lobes 0.964 ± 0.02, parotid glands 0.919 ± 0.03. Dose validation results were summarized in Table 1.

Conclusion:

Using a latent diffusion model, this model-assisted protocol enables ART trigger monitoring with high image fidelity (MAE 9.0 HU, SSIM 0.973) and dosimetric accuracy (<2.5% error), reducing ART-week scanning frequency by 40%.

Table 1. Absolute errors of dose-volume metrics between predicted and real FBCT images

Note: The right three metrics refer to planning target volumes with expanded or contracted margins; V denotes the percentage volume receiving the specified dose (CTV1 and CTV2: 95% of prescription dose).

Dose - Volume Metrics

Absolute Error (cGy)

Relative Error (%)

Dose - Volume Metrics

Absolute Error (%)

Brainstem_D1%

118.2 ± 72.6

2.2 ± 1.5

T_PGTVp_V6996cGy

0.1 ± 0.1

Temporal Lobes_D1%

62.3 ± 55.9

1.0 ± 0.9

T_PCTV1_V5706cGy

0

Parotids_MeanDose

90.0 ± 71.9

2.4 ± 1.9

T_PCTV2_V5141cGy

0.4 ± 0.4