Main Session
Sep
28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology
2494 - Quantitative Validation of AI Platforms for Clinical Evidence Retrieval: Gold-Standard Comparison and Scoring Rubric Analysis of LDRT Trials in Osteoarthritis
Presenter(s)
Aishwarya Kalluri, - Orlando College of Osteopathic Medicine, Winter Garden, FL
A. Kalluri1, K. De La Cruz Quezada1, N. Persaud1, J. M. Rineer2, and T. Dvorak2; 1Orlando College of Osteopathic Medicine, Winter Garden, FL, 2Orlando Health Cancer Institute, Orlando, FL
Purpose/Objective(s):
Artificial intelligence (AI) platforms are increasingly used to retrieve medical evidence, but concerns regarding accuracy, reproducibility, and hallucinated outputs warrant validation. Five large language model platforms were quantitatively evaluated for reliability in identifying randomized controlled trials investigating low-dose radiation therapy (LDRT) for osteoarthritis using a predefined gold-standard dataset and structured scoring rubric.Materials/Methods:
Five AI platforms were evaluated using an identical structured prompt, with 10 independent runs per platform. Outputs were assessed against a manually curated gold-standard set of 5 key randomized controlled trials on low-dose radiation therapy (LDRT) for osteoarthritis using a predefined scoring rubric across five domains: Study Identification (0–4), Accuracy (0–2), Completeness (0–2), Consistency (0–2), and Hallucination (0–2), generating a composite score (best 12) per run. Platforms included ChatGPT-4.0, Gemini-2.0, Meta AI Llama, Perplexity AI Sonar, and Claude Sonnet 4.0. Composite scores were summarized descriptively and compared using one-way ANOVA with eta-squared effect size.Results:
Mean composite scores (±SD) were: ChatGPT-4.0, 4.3±2.3; Gemini-2.0, 3.8±1.3; Meta AI Llama, 2.6±0.7; Perplexity AI Sonar, 4.0±1.8; Claude Sonnet 4.0, 4.6±2.4. %). Coefficients of variation across platforms ranged from 27% to 54%. Although Claude Sonnet 4.0 demonstrated the highest mean score, it also showed the greatest inter-run variability (coefficient of variation 54%). Hallucinated or unverifiable trial data occurred in 100% of runs across all platforms, with severe fabrication observed in up to 80% of outputs for one platform. One-way ANOVA demonstrated no statistically significant difference in composite scores (F=1.54, p=0.21).Conclusion:
No platform demonstrated adequate reliability, with the highest mean composite score reaching only 4.6/12. Importantly, the best-performing model also exhibited the greatest run-to-run variability (54%), highlighting a fundamental lack of reproducibility even under identical prompting conditions. Together with universal hallucination across all platforms, these findings indicate that current AI tools remain too inaccurate and inconsistent for clinical, research, or patient use. Such variability is incompatible with clinical and research workflows, reinforcing the need for rigorous human oversight before AI-generated outputs are used to support evidence synthesis or clinical decision-making.