Main Session
Sep 29
PQA 05 - Physics

2987 - Benchmarking Open-Source Foundation Model (MedSAM2) Against Commercial Auto-Segmentation for Head and Neck OARs and Target Volumes on Simulation CT and Enhanced CBCT

12:30pm - 01:45pm ET
Poster Hall - Exhibit Hall A
Screen: 32
POSTER

Presenter(s)

Omar Awad, MD, MS Headshot
Omar Awad, MD, MS - Baylor College of Medicine, Houston, TX

R. Dotel1, B. Sun2, Y. Han2, S. A. Zaid2, P. Pathak2, C. Stewart1,2, C. Claunch2, O. Awad2, and A. S. Mohamed2; 1Baylor College of Medicine, Houston, TX, 2Department of Radiation Oncology, Dan L. Duncan Comprehensive Cancer Center, Baylor College of Medicine, Houston, TX

Purpose/Objective(s):

To benchmark MedSAM2, an open-source foundation model, against three commercial auto-segmentation platforms: Radformation, RayStation Deep Learning (DL), and RayStation Model-Based Segmentation (MBS) for head-and-neck auto-segmentation on simulation CT (CTSim) and enhanced CBCT.

Materials/Methods: Five head-and-neck cancer patients undergoing definitive RT with paired Sim CT and enhanced CBCT were included. Ground truth was STAPLE consensus from 6 independent Radiation Oncology Specialists. MedSAM2 was evaluated on 12 regions of interest (ROIs); commercial tools on their available templates (2-5 structures each). Five representative slices per structure per patient were assessed using DSC and HD95. Analyses were exploratory given n=5.

Results: MedSAM2 achieved DSC =0.85 for 9/12 ROIs on Sim CT, including GTVp (0.883), Submandibular Gland (0.943), Oral Cavity (0.915), Masseter (0.885), Soft Palate (0.891), and Genioglossus (0.894). Mandible DSC on Sim CT was poor for standard-resolution patients (median 0.188, range 0.068–0.349; n=4, pixel spacing ~1.17mm) but acceptable for the high-resolution patient (DSC=0.866, pixel spacing 0.488mm); CBCT performance was excellent across all patients (median DSC=0.955). GTVp was stable across modalities (Sim: 0.883, CBCT: 0.887, p=0.438). No significant Sim-to-CBCT degradation was observed for any tool (all p>0.05). While MedSAM2 matched RayStation DL in Parotid DSC (Sim: 0.846 vs 0.898, p=0.125) and exceeded it for Submandibular DSC (Sim: 0.926 vs 0.904; CBCT: 0.924 vs 0.856), Parotid HD95 was substantially higher (Sim: 9.37 vs 1.17mm; CBCT: 17.60 vs 1.02mm, p=0.062), indicating imprecise boundary delineation despite comparable volumetric overlap. Across evaluated ROIs, mean DSC was: MedSAM2 (Sim: 0.856, CBCT: 0.856), RayStation DL (Sim: 0.899, CBCT: 0.887), Radformation (Sim: 0.801, CBCT: 0.838), and RayStation MBS (Sim: 0.704, CBCT: 0.728)

Conclusion: MedSAM2 demonstrated competitive volumetric overlap across most ROIs compared to any single commercial tool with no significant Sim-to-CBCT degradation, supporting the feasibility of CBCT-based adaptive auto-segmentation. Elevated Parotid HD95 warrants attention for boundary-sensitive planning. Critically, MedSAM2's current slice-by-slice prompting requirement limits throughput; volumetric automation or efficient prompt propagation strategies are essential before clinical integration. Open-source foundation models represent viable complements to commercial tools for institutions with limited template coverage, pending validation in larger, multi-institutional cohorts