Main Session
Sep
30
QP 46 - Imaging for Planning
1275 - Robust and Precise 4D-CT Ventilation Imaging via Spatially-Encoded Attention U-Net: A Multi-Modality Validation for Functional Lung Avoidance
Presenter(s)
Nannan Qin, MS - The First Affiliated Hospital of Bengbu Medical University, Bengbu, Anhui
N. Qin1, H. Cai1, J. Wang1, S. Song1, Y. Zhou1, and A. Wu2; 1Department of Radiation Oncology, The First Affiliated Hospital of Bengbu Medical University, Bengbu, Anhui, China, 2Department of Radiation Oncology, The First Affiliated Hospital of University of Science and Technology of China, HeFei, Anhui, China
Purpose/Objective(s):
CT-based ventilation imaging (CTVI) provides a cost-effective alternative to PET/SPECT for functional lung avoidance radiation therapy (RT). However, current deep learning (DL) methods often struggle with generalization across different validation modalities (e.g., PET vs. SPECT) and lack spatial awareness of gravity-dependent ventilation gradients. We propose a novel DL framework integrating explicit spatial encoding and attention mechanisms to generate robust, high-fidelity ventilation maps from standard 4D-CTsMaterials/Methods:
Forty-six lung cancer patients from the VAMPIRE challenge dataset were utilized, including 25 with Galligas-PET and 21 with DTPA-SPECT ventilation ground truths (GT). We developed a 3D Attention U-Net architecture taking a 6-channel input: End-Inhale/Exhale CTs, a temporal difference map to capture motion, and explicit X-Y-Z spatial coordinate maps to model physiological gravity gradients. A hybrid loss function incorporating a weighted Dice term was designed to prioritize the localization of high-ventilation regions. Model performance was rigorously evaluated using 5-fold cross-validation. Metrics included voxel-wise Spearman’s rank correlation coefficient (rs) and Dice similarity coefficient (DSC) for the top-20% functional volumes.Results:
Quantitative results are summarized in Table 1. The proposed method achieved a state-of-the-art overall mean rs of 0.66 ± 0.11, significantly outperforming the VAMPIRE benchmark (traditional deformable registration, mean rs˜0.45). Notably, the model demonstrated exceptional cross-modality robustness, yielding consistent accuracy for both high-quality PET (rs=0.66) and lower-resolution SPECT (rs=0.66) cohorts. The segmentation of high-functional lung regions achieved a mean DSC of 0.51 ± 0.10, with superior performance in identifying ventilation defects (DSCLow=0.58). Qualitative analysis confirmed that the spatial encoding successfully captured physiological gravity-dependent ventilation gradients, even in cases with noisy GTConclusion:
Our spatially-encoded attention network generates robust and physiologically plausible ventilation images from 4D-CT, effectively overcoming the domain shift between CT and nuclear medicine modalities. The high correlation and accurate functional localization suggest that this method is a viable, radiation-free tool for widespread implementation of functional avoidance RT planning. Table 1:Quantitative comparison of the proposed method against traditional benchmarks across PET and SPECT modalities (5-Fold Cross-Validation Results) *BM-DIR values derived from the original VAMPIRE challenge summary (Med. Phys. 2019).| Metric\Modality | Overall (N=46) | Galligas PET(n=25) | DTPA-SPECT(n=25) |
| Spearman Correlation (rs) | |||
| BM-DIR* | 0.53± 0.10 | 0.49 ± 0.16 | |
| Proposed [Ours] | 0.66±0.11 | 0.66 ± 0.10 | 0.66 ± 0.12 |
| Dice Coefficient (Top 20%) | |||
| BM-DIR* | 0.47 ± 07 | 0.45± 0.11 | |
| Proposed [Ours] | 0.51±0.10 | 0.51 ± 0.10 | 0.50 ± 0.10 |