2600 - Cross-Platform Evaluation of Brain Stereotactic Radiosurgery Plan Quality Using a Dual-Model Deep Learning Framework
R. Shah1, Z. Xiao2, W. Cao2, G. O'Neill1, M. Bloomfield1, A. Alluri3, W. Shi2, Y. Chen2, H. Liu2, and W. Wang2; 1Sidney Kimmel Medical College at Thomas Jefferson University, Philadelphia, PA, 2Department of Radiation Oncology, Sidney Kimmel Medical College & Cancer Center at Thomas Jefferson University, Philadelphia, PA, 3Drexel University, Philadephia, PA
Purpose/Objective(s): Despite known dosimetric differences among SRS delivery platforms, objective cross-platform plan quality comparisons are constrained by the lack of a standardized reference dose distribution. We tested whether a dual-model deep learning framework could generate benchmark dose distributions to quantify systematic differences in dose falloff among three platforms: Gamma Knife (GK), CyberKnife (CK), and LINAC.
Materials/Methods:
We retrospectively evaluated 64 intracranial targets from 30 patients treated with GK (10 patients, 10 targets), CK with fixed cone collimation (CK-Cone; 6 patients, 10 targets), CK with multileaf collimator (CK-MLC; 4 patients, 12 targets) and LINAC (10 patients, 32 targets). Target volumes ranged from 0.1 to 43.3 cc (median 0.9 cc). For each target, an “ideal” benchmark dose distribution was predicted by an in-house deep learning framework. This framework employed two independent U-Net models: one trained on small targets (<1 cc) and one on large targets (=1 cc), designed to predict optimal dose falloff. Clinical plans were quantitatively evaluated against these benchmark predictions using the Brain Sparing Index (BSI) and V12Gy difference.Results: Across 64 targets, the mean BSI was 0.63 ± 0.23, indicating overall higher peri-target dose spillage in the clinical plans relative to the deep learning-predicted benchmark. GK (BSI = 0.85/0.56 for small/large targets) achieved superior dose falloff for small targets while struggling with large targets compared to LINAC (BSI = 0.60/0.70 for small/large targets). For CK, fixed cone collimation (BSI = 0.58/0.51 for small/large targets) was more suitable for small targets, and MLC (BSI = 0.48/0.75 for small/large targets) achieved better dose falloff for large targets. At the patient level, there was higher mean V12Gy in clinical plans (13.2 ± 20.5 cc) relative to the benchmark (8.9 ± 14.2 cc).
Conclusion: The size-stratified deep learning framework enables standardized, cross-platform evaluation of SRS dose falloff. Clinical plans demonstrated modality-associated differences relative to the benchmark. This framework provides an objective, reproducible tool for benchmarking SRS plan quality and selecting an evidence-based delivery platform.
Table 1. Clinical Plan Agreement with Dual-Model Deep Learning-Predicted Benchmark for Brain SRS.| Modality | GK | CK-Cone | CK-MLC | LINAC | All |
| Patients | 10 | 6 | 4 | 10 | 30 |
| # of Total Targets | 10 | 10 | 12 | 32 | 64 |
| All Targets BSI (mean ± SD) | 0.64 ± 0.22 | 0.55 ± 0.15 | 0.66 ± 0.33 | 0.64 ± 0.19 | 0.63 ± 0.23 |
| # of Small Targets (< 1 cc) | 3 | 6 | 4 | 21 | 34 |
| Small Targets BSI (mean ± SD) | 0.85 ± 0.28 | 0.58 ± 0.18 | 0.48 ± 0.10 | 0.60 ± 0.21 | 0.61 ± 0.22 |
| # of Large Targets (= 1 cc) | 7 | 4 | 8 | 11 | 30 |
| Large Targets BSI (mean ± SD) | 0.56 ± 0.09 | 0.51 ± 0.08 | 0.75 ± 0.37 | 0.70 ± 0.13 | 0.66 ± 0.23 |
| Benchmark V12Gy (cc, mean ± SD) | 1.5 ± 1.4 | 6.8 ± 7.8 | 37.3 ± 16.7 | 6.4 ± 8.2 | 8.9 ± 14.2 |
| Clinical V12Gy (cc, mean ± SD) | 2.6 ± 2.3 | 11.3 ± 12.3 | 51.9 ± 28.9 | 9.4 ± 9.9 | 13.2 ± 20.5 |