Main Session
Sep 29
PQA 05 - Physics

3144 - Preserving Boundary Accuracy in Compact 3D nnU-Net via Boundary-Privileged Knowledge Distillation

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 17
POSTER

Presenter(s)

Zenan Sun, MS - University of Texas Health Science Center at Houston, Houston, TX

Z. Sun1, Q. Lan1, and J. Duan2; 1UTHealth Houston, Houston, TX, 2The University of Texas MD Anderson Cancer Center, Houston, TX

Purpose/Objective(s): Knowledge distillation enables fast 3D segmentation when full-scale models are impractical in resource and time-sensitive clinical workflows such as adaptive radiotherapy. However, conventional KD overemphasizes background and global similarity, degrading boundary and small-structure accuracy, that is critical in radiotherapy where millimeter-level contour boundary misses can materially change DVH-driven optimization under steep dose gradients and tight OAR constraints. To resolve this, we propose Boundary Saliency Injected Decoupled (BSID) Distillation, a distillation strategy that emphasizes boundary learning.

Materials/Methods: Our method first derives a teacher-based boundary importance map using a fixed 3D Prewitt edge operator, then enforces boundary-gradient alignment between teacher and student, and finally uses this map to weight only the target/ distillation term while keeping the background term global. We evaluated an nnU-Net teacher and compact student model under a fixed architecture. The student was trained with no distillation, logit-based KD, structured knowledge distillation (SKD), and intra-class feature variation distillation (IFVD), and our method, and tested on LiTS 2017 (liver tumor, n = 131) and KiTS 2023 (kidney tumor, n = 599).

Results: Table 1 summarizes the performance gains of student models with different distillation strategies. Compared with the supervised student baseline, our method improves LiTS 2017 DSC/NSD(mm)/HD95(mm) from 70.67/62.62/70.60 to 78.78/72.31/47.80, and improves KiTS 2023 from 68.65/57.72/168.44 to 84.61/74.52/35.07, outperforming logit-KD, SKD, and IFVD, which show limited or even negative gains on KiTS. Notably, these improvements are achieved without increasing model complexity: both proposed and baseline model uses ~75% fewer parameters than the teacher model.

Conclusion: BSID improves boundary fidelity and overall segmentation performance in compact 3D nnU-Net students without adding inference-time overhead, supporting efficient deployment in boundary-sensitive, time-critical clinical workflows such as online adaptive radiotherapy. Moreover, BSID operates entirely in output space as architecture-agnostic and can be applied to other 3D segmentation backbones.

Table 1. Tumor segmentation performance avg(std) on LiTS 2017 and KiTS 2023

Method

LiTS 2017

KiTS 2023

Dice?

NSD(mm) ?

HD95(mm) ?

Dice?

NSD(mm) ?

HD95(mm) ?

Teacher

83.21(2.03)

79.18(2.74)

15.26(3.12)

88.87(0.92)

81.98(1.14)

27.13(5.77)

Student(Baseline)

70.67(2.58)

62.62(2.88)

70.60(8.66)

68.65(2.09)

57.72(1.60)

168.44(11.03)

logits

74.47(2.49)

68.41(2.80)

65.98(9.91)

58.42(2.26)

47.74(1.56)

192.02(9.37)

skd

70.68(2.69)

62.77(2.89)

78.88(8.41)

61.85(2.35)

51.95(1.71)

176.02(10.34)

ifvd

74.67(2.67)

68.01(2.90)

71.51(9.22)

65.39(2.24)

55.20(1.70)

173.46(10.67)

Ours

78.78(2.29)

72.31(2.89)

47.80(7.82)

84.61(1.19)

74.52(1.28)

35.07(6.81)