Main Session
Sep 28
QP 13 - Pixels to Plans: AI-Driven Contouring and Imaging Innovation

1073 - Eliminating Uncertainties in Post-Treatment Breast Cosmesis Evaluation using Vision-Language Artificial Intelligence

03:10pm - 03:15pm ET
Room 162

Presenter(s)

Sangjoon Park, MD, PhD - Yonsei University College of Medicine, Seoul, Seoul

S. Park1, H. K. Byun2, Y. Cho II3, J. Y. Kim4, D. W. Lee5, S. Y. Song5, and J. Chang1; 1Department of Radiation Oncology, Yonsei Cancer Center, Heavy Ion Therapy Research Institute, Yonsei University College of Medicine, Seoul, Korea, Republic of (South), 2Department of Radiation Oncology Yongin Severance Hospital, Yonsei University College of Medicine, Yongin, Korea, Republic of (South), 3Department of Radiation Oncology, Gangnam Severance Hospital, Yonsei University College of Medicine, Seoul, Korea, Republic of (South), 43Department of Surgery, Yonsei University College of Medicine, Seoul, Korea, Republic of (South), 5Department of Plastic and Reconstructive Surgery, Severance Hospital, Yonsei University College of Medicine, Seoul, Korea, Republic of (South)

Purpose/Objective(s): Aesthetic outcomes following breast-conserving therapy are critical determinants of quality of life, yet current evaluation methods are limited by the subjectivity of existing scales and substantial inter-observer variability. This study investigated the feasibility of Large Vision-Language Models (LVLMs) as an objective, label-independent framework for breast cosmesis assessment and evaluated their potential to provide consistent evaluations free from exposure-dependent learning effects.

Materials/Methods: We analyzed 89 frontal-view breast images from the RTOG 1014 trial, with expert-consensus labels established by consensus grading from six radiation oncologists using the Global Cosmetic Score. Multiple LVLMs, including the GPT-5 family, Gemini 2.0 Flash, and Claude-4.5 Haiku, were evaluated after systematic optimization of prompt complexity, scoring scales, and few-shot strategies. A multidisciplinary expert survey involving seven clinicians was conducted on a subset of 34 randomly selected cases, and learning effects were assessed through sequential subgroup analysis.

Results: Optimal performance was achieved with multi-metric evaluation, detailed prompts, continuous scoring, and one-shot exemplars. The optimized GPT-5-based framework achieved significant correlation with expert consensus (Spearman's ? = 0.697; 95% CI: 0.564–0.789) and excellent reproducibility (ICC = 0.982; 95% CI: 0.974–0.987). High concordance was observed for macroscopic features such as volume, symmetry, and contour, whereas performance was lower for subtle changes including scar appearance. While human evaluators achieved higher overall agreement with labels, they demonstrated a clear learning effect, with lower concordance in early evaluations that progressively improved with cumulative case exposure. In contrast, LVLM-based assessments remained consistent regardless of case number.

Conclusion: LVLMs offer a robust and reproducible method for objectively evaluating breast cosmesis without task-specific training. By providing consistent assessments independent of exposure-dependent learning, this approach more closely approximates real-world clinical practice where cosmesis is evaluated on a per-patient basis. This AI-driven framework may serve as a standardized calibration tool for clinical decision-making and research endpoints in breast oncology.