Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2433 - Quantitative Assessment of AI Platforms in Extracting and Synthesizing Clinical Evidence from Retrospective Brachytherapy Studies for Primary Vaginal Cancer

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 15

Presenter(s)

Tomas Dvorak, MD - Orlando Health Cancer Institute, Orlando, FL

K. De La Cruz Quezada1, A. Kalluri1, N. Persaud1, and T. Dvorak2; 1Orlando College of Osteopathic Medicine, Winter Garden, FL, 2Orlando Health Cancer Institute, Orlando, FL

Purpose/Objective(s):

The use of artificial intelligence (AI) platforms to support clinical decision making and treatment planning is rapidly increasing in oncology; however, their reliability for identifying and synthesizing clinical evidence, particularly in rare malignancies and uncommon therapeutic contexts, remains insufficiently characterized. This study quantitatively evaluates the performance of multiple AI platforms in extracting metadata from retrospective studies of low-dose-rate brachytherapy for primary vaginal cancer, a rare malignancy accounting for <1% of gynecologic cancers.

Materials/Methods:

Five AI platforms were evaluated using an identical structured prompt, with 10 independent runs per platform. Outputs were assessed against a manually curated gold-standard set of 5 key retrospective studies using a predefined scoring rubric across five domains: Study Identification (0–5), Accuracy (0–2), Completeness (0–2), Consistency (0–2), and Hallucination (0–2), generating a composite score (maximum 13) per run. Platforms included ChatGPT 4.0, Gemini Flash 3.0, Microsoft Copilot Smart 5.0, Perplexity AI Sonar, and Meta AI Llama 4.0. Composite scores were summarized descriptively and compared using one-way ANOVA with eta-squared effect size.

Results:

Mean composite scores ranged from 2.0 to 5.5. Meta AI Llama 4.0 demonstrated the highest mean performance (5.5 ± 0.85), followed by ChatGPT 4.0 (4.3 ± 1.34) and Microsoft Copilot Smart 5.0 (4.0 ± 2.28), while Gemini Flash 3.0 showed lower but highly stable performance (2.9 ± 0.32) and Perplexity AI Sonar exhibited the lowest mean score with marked instability (2.0 ± 2.26). Stability varied substantially, with coefficients of variation of 15%, 31%, 57%, 11%, and 113%, respectively. One-way ANOVA demonstrated statistically significant differences among platforms (F = 7.63, p < 0.001) with a large effect size (?² = 0.40); no platform achieved complete identification of the manually curated gold-standard set. Hallucination rates were observed in 100%, 40%, 60%, 100%, and 90% of the 10 runs, respectively.

Conclusion: AI platform choice significantly influenced performance in extracting and synthesizing retrospective brachytherapy evidence. Despite measurable differences in composite scores, all platforms demonstrated incomplete study identification and frequent hallucinations, underscoring the critical need for careful human oversight before current AI-generated outputs can be reliably used for oncologic evidence synthesis or to inform clinical decision making, particularly in rare malignancies and uncommon treatment contexts.