Main Session
Sep 28
PQA 04 - Breast Cancer, Patient Reported Outcomes/QoL/Survivorship, Functional Radiation Medicine, Hematologic Malignancies, Palliative Care, and International/Global Oncology

2697 - A Blinded Comparison of Artificial Intelligence-Generated vs. Physician-Drawn Target Contours for Partial Breast Irradiation

03:00pm - 04:00pm ET
Poster Hall - Exhibit Hall A
Screen: 6
POSTER

Presenter(s)

Eva Berlin, MD Headshot
Eva Berlin, MD - Hospital of the University of Pennsylvania, Philadelphia, PA

E. Berlin, B. Byrd, N. Molina, S. Philbrook, M. Kassick, G. M. Freedman, N. K. Taunk, W. R. Green, and R. McBeth; Department of Radiation Oncology, University of Pennsylvania, Philadelphia, PA

Purpose/Objective(s):

We hypothesized that artificial intelligence (AI)-generated partial breast irradiation (PBI) target contours would achieve comparable clinical quality to the gold standard of manual physician-drawn contours, as assessed by blinded review.

Materials/Methods:

This single-institution study was conducted under an IRB-approved protocol. An in-house deep learning platform (nnU-Net) was trained on 533 GTV contours from treated PBI cases drawn by 5 physicians. CTV and PTV structures were derived via rule-based expansions per institutional protocol based on the Florence Trial. AI contours were generated on 25 consecutively treated (September-November 2025) non-training PBI cases. Bilateral disease, reconstruction, and incomplete or incorrectly named contours were excluded. Three academic breast radiation oncologists independently performed a blinded review of AI versus manual contours presented in random order. GTV, CTV, and PTV quality were scored: 1=major revisions, 2=minor revisions, 3=no changes. Scores =2 indicated clinically acceptable contours. Reviewers also selected overall contour preference for each of the 25 cases (A, B, or no preference). A pre-specified one-sided binomial test assessed contour preference. Paired Wilcoxon signed-rank tests compared mean contour quality scores. Dice coefficients quantified geometric agreement between AI and manual contours. Inter-rater reliability was assessed using Gwet's AC1 to address the prevalence paradox.

Results:

Three reviewers evaluated 25 PBI cases across 3 target structures (GTV, CTV, PTV) resulting in 75 case reviews and a total of 225 contour assessments. Inter-rater reliability was substantial (AC1=0.65; pairwise agreement 70.7%). AI contours were scored as clinically acceptable (score =2 in all 3 structures) in 85.3% of cases vs 77.3% for manual contours. When contour preference was indicated (in 45/75 case reviews), AI contours were preferred (29/45, 64%) vs manual (16/45, 36%), p=0.04. Paired comparisons of mean contour quality scores showed no significant differences between AI and manual contours for any structure (Table). Dice coefficients demonstrated good agreement: GTV 0.69±0.22, CTV 0.83±0.12, PTV 0.83±0.10; CTV/PTV values reflect GTV agreement plus shared expansion rules.

Conclusion:

In blinded review by breast radiation oncology specialists, AI-generated PBI target contours demonstrated comparable quality to manual physician-drawn contours, with no significant differences in mean quality scores and substantial inter-rater agreement. AI contours were significantly preferred over manual contours. These findings support the integration of AI-generated target contours into PBI treatment planning to improve consistency and efficiency while maintaining clinical quality.

AI tool Claude Opus, Anthropic was used for language refinement. Authors are responsible for scientific content.

Structure

Manual

AI

P-value

GTV

2.59

2.64

0.68

CTV

2.77

2.77

1.00

PTV

2.64

2.75

0.40