Main Session
Sep 28
PQA 04 - Breast Cancer, Patient Reported Outcomes/QoL/Survivorship, Functional Radiation Medicine, Hematologic Malignancies, Palliative Care, and International/Global Oncology

2747 - Matching Clinical Contouring Performance: A Blinded-Rater Study of a Novel Transformer-Based AI Model (TUM-SAM) for Hepatocellular Carcinoma (HCC) Contouring in Mongolia

03:00pm - 04:00pm ET
Poster Hall - Exhibit Hall A
Screen: 20
POSTER

Presenter(s)

Baozhou Sun, PhD, MBA Headshot
Baozhou Sun, PhD, MBA - Baylor College of Medicine, Houston, Tx

Y. Han1, A. N. Hanania1, S. S. Desai1, A. Anderson2, P. Pathak1, D. A. Hamstra1, S. A. Zaid1, M. Minjgee3, E. Vanchinbazar3, B. Batsuuri3, U. Tsegmid3, B. Magsar3, U. Bayasgalan4, and B. Sun1; 1Department of Radiation Oncology, Dan L. Duncan Comprehensive Cancer Center, Baylor College of Medicine, Houston, TX, 2Fred & Pamela Buffett Cancer Center - Nebraska Medical Center, Omaha, NE, 3National Cancer Center of Mongolia, Ulaanbaatar, Mongolia, 4National Cancer Center Korea, Goyang, Korea, Republic of (South)

Purpose/Objective(s):

Mongolia has the world’s highest incidence of HCC. Quality definitive radiation oncology care remains limited by MRI availability and subspecialty contouring expertise. In this setting, contrast-enhanced CT (CECT) tumor contouring is essential but challenged by substantial inter-observer variability. We developed an AI transformer-based model (TUM-SAM) for CECT-based segmentation. We hypothesized that in a blinded head-to-head evaluation by US expert clinicians, preference rates would not differ between AI-generated contours and those generated by attending radiation oncologists in Mongolia, and that AI contouring would reduce contouring time.

Materials/Methods:

Three Mongolian radiation oncologists (>7 years’ experience) independently contoured gross tumor volumes (GTVs) for 10 patients. A multi-expert consensus GTV reference was established by three US-based, board-certified gastrointestinal radiation oncologists. AI contours were generated using TUM-SAM trained on multi-institutional US CECT datasets. Quantitative agreement was assessed using Dice Similarity Coefficient (DSC). The primary endpoint was clinician preference in blinded head-to-head comparisons, analyzed using a Bradley–Terry model to estimate win probabilities while accounting for rater bias and case complexity. Manual and AI-assisted contouring times were recorded.

Results:

Relative to the US multi-expert consensus reference, Mongolian physicians’ contours achieved mean DSCs of 0.78 (95% CI 72% – 85%), 0.83 (95% CI 83% – 84%), and 0.84 (95% CI 76% – 92%), while TUM-SAM achieved a mean DSC of 0.75 (95% CI: 0.70–0.80). Meanwhile, in blinded head-to-head evaluations, no significant difference in clinician preference was observed between AI-generated and physician-generated contours. Bradley–Terry analysis yielded AI win probabilities of 49.6% (95% CI: 46.6–53.1%) against Physician A, 46.9% (95% CI: 43.0–50.9%) against Physician B, and 48.0% (95% CI: 46.0–50.0%) against Physician C, all of which demonstrated no statistically significant preference difference. TUM-SAM generated contours in less than 1 minute per case, whereas manual contouring by Mongolian physicians required a mean of 20 minutes per case.

Conclusion:

Although Mongolian physicians’ contours demonstrated higher geometric agreement with the US consensus reference, TUM-SAM contours had no statistically significant difference in blinded clinician preference and provided clinically-relevant reductions in contouring time. These findings suggest AI-based HCC contouring may be clinically acceptable and serve as a scalable tool to enhance efficiency and promote consistent contouring standards in resource-constrained settings.