2987 - Benchmarking Open-Source Foundation Model (MedSAM2) Against Commercial Auto-Segmentation for Head and Neck OARs and Target Volumes on Simulation CT and Enhanced CBCT
Presenter(s)
R. Dotel1, B. Sun2, Y. Han2, S. A. Zaid2, P. Pathak2, C. Stewart1,2, C. Claunch2, O. Awad2, and A. S. Mohamed2; 1Baylor College of Medicine, Houston, TX, 2Department of Radiation Oncology, Dan L. Duncan Comprehensive Cancer Center, Baylor College of Medicine, Houston, TX
Purpose/Objective(s):
To benchmark MedSAM2, an open-source foundation model, against three commercial auto-segmentation platforms: Radformation, RayStation Deep Learning (DL), and RayStation Model-Based Segmentation (MBS) for head-and-neck auto-segmentation on simulation CT (CTSim) and enhanced CBCT.Materials/Methods: Five head-and-neck cancer patients undergoing definitive RT with paired Sim CT and enhanced CBCT were included. Ground truth was STAPLE consensus from 6 independent Radiation Oncology Specialists. MedSAM2 was evaluated on 12 regions of interest (ROIs); commercial tools on their available templates (2-5 structures each). Five representative slices per structure per patient were assessed using DSC and HD95. Analyses were exploratory given n=5.
Results: MedSAM2 achieved DSC =0.85 for 9/12 ROIs on Sim CT, including GTVp (0.883), Submandibular Gland (0.943), Oral Cavity (0.915), Masseter (0.885), Soft Palate (0.891), and Genioglossus (0.894). Mandible DSC on Sim CT was poor for standard-resolution patients (median 0.188, range 0.068–0.349; n=4, pixel spacing ~1.17mm) but acceptable for the high-resolution patient (DSC=0.866, pixel spacing 0.488mm); CBCT performance was excellent across all patients (median DSC=0.955). GTVp was stable across modalities (Sim: 0.883, CBCT: 0.887, p=0.438). No significant Sim-to-CBCT degradation was observed for any tool (all p>0.05). While MedSAM2 matched RayStation DL in Parotid DSC (Sim: 0.846 vs 0.898, p=0.125) and exceeded it for Submandibular DSC (Sim: 0.926 vs 0.904; CBCT: 0.924 vs 0.856), Parotid HD95 was substantially higher (Sim: 9.37 vs 1.17mm; CBCT: 17.60 vs 1.02mm, p=0.062), indicating imprecise boundary delineation despite comparable volumetric overlap. Across evaluated ROIs, mean DSC was: MedSAM2 (Sim: 0.856, CBCT: 0.856), RayStation DL (Sim: 0.899, CBCT: 0.887), Radformation (Sim: 0.801, CBCT: 0.838), and RayStation MBS (Sim: 0.704, CBCT: 0.728)
Conclusion: MedSAM2 demonstrated competitive volumetric overlap across most ROIs compared to any single commercial tool with no significant Sim-to-CBCT degradation, supporting the feasibility of CBCT-based adaptive auto-segmentation. Elevated Parotid HD95 warrants attention for boundary-sensitive planning. Critically, MedSAM2's current slice-by-slice prompting requirement limits throughput; volumetric automation or efficient prompt propagation strategies are essential before clinical integration. Open-source foundation models represent viable complements to commercial tools for institutions with limited template coverage, pending validation in larger, multi-institutional cohorts