Main Session
Sep 29
PQA 05 - Physics

3213 - Learning What Can Move: Patient-Specific Motion Subspaces from a Single Scan

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 24
POSTER

Presenter(s)

Yinheng Zhu, PhD - Department of Radiation Oncology, UT Southwestern Medical Center, Dallas, TX

Y. Zhu1, M. Chen1, Q. Wang1, X. Gu2, and W. Lu3; 1Medical Artificial Intelligence and Automation (MAIA) Lab, Department of Radiation Oncology, UT Southwestern Medical Center, Dallas, TX, 2Stanford University Department of Radiation Oncology, Palo Alto, CA, 3Department of Radiation Oncology, University of Texas Southwestern Medical Center, Dallas, TX

Purpose/Objective(s): Motion estimation naturally requires two time points, making it inherently a relational quantity. This constrains how motion can be learned during training, as the entanglement of paired data makes cross-modality and cross-dimensional scenarios particularly challenging. We approach this differently by modeling the motion subspace rather than motion itself. Unlike motion, the motion subspace is an intrinsic property potentially inferable from a single observation, since anatomical configuration inherently constrains the space of plausible deformations. We investigate whether a learning system can infer patient-specific motion modes from a single static scan for unified multi-modality motion reconstruction.

Materials/Methods: We parameterize deformation as K scalar spatial motion modes, capturing where motion (co)-occurs, multiplied by global vector coefficients, specifying 3D direction and magnitude. A neural network predicts patient-specific motion modes from the fixed image alone. A key challenge in learning such basis-coefficient combination is the inherent scale ambiguity, which causes unstable or failure training. We address this with a differentiable closed-form coefficient solver that, given predicted modes and any observation, uniquely determines the optimal coefficients via linearized least-squares.

Results: The learned motion modes effectively capture patient-specific motion subspace that consistently generalizes across all other phases. With K=2 modes, the method achieved Dice of 0.921, HD95 of 2.926 mm, and PSNR of 33.06 dB in the 3D-3D setting. A baseline using direct coefficient regression instead of the closed-form solver performed substantially worse (Dice 0.882, HD95 4.283 mm). The same learned modes, without retraining, supported four observation modalities with Dice above 0.90: 2D image, 2D mask, and surface points. Once modes are pre-computed, the closed-form coefficient solver requires no network forward pass and completes in under 10 ms. The predicted modes exhibit interpretable spatial patterns, with co-moving structures sharing similar activations.

Conclusion: This work demonstrates that patient-specific motion modes can be inferred from static anatomy alone. The learned subspace can be used directly or seamlessly integrated into conventional DIR frameworks as learned regularization. The asymmetric training, where only the fixed image is input to the network, offers potential for cross-modality applications. The compact parameterization makes the approach well-suited for scenarios with sparse observations.