Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2600 - Cross-Platform Evaluation of Brain Stereotactic Radiosurgery Plan Quality Using a Dual-Model Deep Learning Framework

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 7
POSTER

R. Shah1, Z. Xiao2, W. Cao2, G. O'Neill1, M. Bloomfield1, A. Alluri3, W. Shi2, Y. Chen2, H. Liu2, and W. Wang2; 1Sidney Kimmel Medical College at Thomas Jefferson University, Philadelphia, PA, 2Department of Radiation Oncology, Sidney Kimmel Medical College & Cancer Center at Thomas Jefferson University, Philadelphia, PA, 3Drexel University, Philadephia, PA

Purpose/Objective(s): Despite known dosimetric differences among SRS delivery platforms, objective cross-platform plan quality comparisons are constrained by the lack of a standardized reference dose distribution. We tested whether a dual-model deep learning framework could generate benchmark dose distributions to quantify systematic differences in dose falloff among three platforms: Gamma Knife (GK), CyberKnife (CK), and LINAC.

Materials/Methods:

We retrospectively evaluated 64 intracranial targets from 30 patients treated with GK (10 patients, 10 targets), CK with fixed cone collimation (CK-Cone; 6 patients, 10 targets), CK with multileaf collimator (CK-MLC; 4 patients, 12 targets) and LINAC (10 patients, 32 targets). Target volumes ranged from 0.1 to 43.3 cc (median 0.9 cc). For each target, an “ideal” benchmark dose distribution was predicted by an in-house deep learning framework. This framework employed two independent U-Net models: one trained on small targets (<1 cc) and one on large targets (=1 cc), designed to predict optimal dose falloff. Clinical plans were quantitatively evaluated against these benchmark predictions using the Brain Sparing Index (BSI) and V12Gy difference.

Results: Across 64 targets, the mean BSI was 0.63 ± 0.23, indicating overall higher peri-target dose spillage in the clinical plans relative to the deep learning-predicted benchmark. GK (BSI = 0.85/0.56 for small/large targets) achieved superior dose falloff for small targets while struggling with large targets compared to LINAC (BSI = 0.60/0.70 for small/large targets). For CK, fixed cone collimation (BSI = 0.58/0.51 for small/large targets) was more suitable for small targets, and MLC (BSI = 0.48/0.75 for small/large targets) achieved better dose falloff for large targets. At the patient level, there was higher mean V12Gy in clinical plans (13.2 ± 20.5 cc) relative to the benchmark (8.9 ± 14.2 cc).

Conclusion: The size-stratified deep learning framework enables standardized, cross-platform evaluation of SRS dose falloff. Clinical plans demonstrated modality-associated differences relative to the benchmark. This framework provides an objective, reproducible tool for benchmarking SRS plan quality and selecting an evidence-based delivery platform.

Table 1. Clinical Plan Agreement with Dual-Model Deep Learning-Predicted Benchmark for Brain SRS.

Modality GK

CK-Cone

CK-MLC

LINAC

All

Patients 10

6

4

10

30

# of Total Targets 10

10

12

32

64

All Targets BSI (mean ± SD) 0.64 ± 0.22

0.55 ± 0.15

0.66 ± 0.33

0.64 ± 0.19

0.63 ± 0.23

# of Small Targets (< 1 cc) 3

6

4

21

34

Small Targets BSI (mean ± SD) 0.85 ± 0.28

0.58 ± 0.18

0.48 ± 0.10

0.60 ± 0.21

0.61 ± 0.22

# of Large Targets (= 1 cc) 7

4

8

11

30

Large Targets BSI (mean ± SD) 0.56 ± 0.09

0.51 ± 0.08

0.75 ± 0.37

0.70 ± 0.13

0.66 ± 0.23

Benchmark V12Gy (cc, mean ± SD) 1.5 ± 1.4

6.8 ± 7.8

37.3 ± 16.7

6.4 ± 8.2

8.9 ± 14.2

Clinical V12Gy (cc, mean ± SD) 2.6 ± 2.3

11.3 ± 12.3

51.9 ± 28.9

9.4 ± 9.9

13.2 ± 20.5