3144 - Preserving Boundary Accuracy in Compact 3D nnU-Net via Boundary-Privileged Knowledge Distillation
Presenter(s)
Z. Sun1, Q. Lan1, and J. Duan2; 1UTHealth Houston, Houston, TX, 2The University of Texas MD Anderson Cancer Center, Houston, TX
Purpose/Objective(s): Knowledge distillation enables fast 3D segmentation when full-scale models are impractical in resource and time-sensitive clinical workflows such as adaptive radiotherapy. However, conventional KD overemphasizes background and global similarity, degrading boundary and small-structure accuracy, that is critical in radiotherapy where millimeter-level contour boundary misses can materially change DVH-driven optimization under steep dose gradients and tight OAR constraints. To resolve this, we propose Boundary Saliency Injected Decoupled (BSID) Distillation, a distillation strategy that emphasizes boundary learning.
Materials/Methods: Our method first derives a teacher-based boundary importance map using a fixed 3D Prewitt edge operator, then enforces boundary-gradient alignment between teacher and student, and finally uses this map to weight only the target/ distillation term while keeping the background term global. We evaluated an nnU-Net teacher and compact student model under a fixed architecture. The student was trained with no distillation, logit-based KD, structured knowledge distillation (SKD), and intra-class feature variation distillation (IFVD), and our method, and tested on LiTS 2017 (liver tumor, n = 131) and KiTS 2023 (kidney tumor, n = 599).
Results: Table 1 summarizes the performance gains of student models with different distillation strategies. Compared with the supervised student baseline, our method improves LiTS 2017 DSC/NSD(mm)/HD95(mm) from 70.67/62.62/70.60 to 78.78/72.31/47.80, and improves KiTS 2023 from 68.65/57.72/168.44 to 84.61/74.52/35.07, outperforming logit-KD, SKD, and IFVD, which show limited or even negative gains on KiTS. Notably, these improvements are achieved without increasing model complexity: both proposed and baseline model uses ~75% fewer parameters than the teacher model.
Conclusion: BSID improves boundary fidelity and overall segmentation performance in compact 3D nnU-Net students without adding inference-time overhead, supporting efficient deployment in boundary-sensitive, time-critical clinical workflows such as online adaptive radiotherapy. Moreover, BSID operates entirely in output space as architecture-agnostic and can be applied to other 3D segmentation backbones.
Table 1. Tumor segmentation performance avg(std) on LiTS 2017 and KiTS 2023| Method | LiTS 2017 | KiTS 2023 | ||||
| Dice? | NSD(mm) ? | HD95(mm) ? | Dice? | NSD(mm) ? | HD95(mm) ? | |
| Teacher | 83.21(2.03) | 79.18(2.74) | 15.26(3.12) | 88.87(0.92) | 81.98(1.14) | 27.13(5.77) |
| Student(Baseline) | 70.67(2.58) | 62.62(2.88) | 70.60(8.66) | 68.65(2.09) | 57.72(1.60) | 168.44(11.03) |
| logits | 74.47(2.49) | 68.41(2.80) | 65.98(9.91) | 58.42(2.26) | 47.74(1.56) | 192.02(9.37) |
| skd | 70.68(2.69) | 62.77(2.89) | 78.88(8.41) | 61.85(2.35) | 51.95(1.71) | 176.02(10.34) |
| ifvd | 74.67(2.67) | 68.01(2.90) | 71.51(9.22) | 65.39(2.24) | 55.20(1.70) | 173.46(10.67) |
| Ours | 78.78(2.29) | 72.31(2.89) | 47.80(7.82) | 84.61(1.19) | 74.52(1.28) | 35.07(6.81) |