2617 - Evaluating Clinical Realism of Synthetic Radiology Reports to Facilitate Development of PREEMPT Trial Screening
Presenter(s)
M. Sridhar1, K. Sethia1, J. Kang2, M. Nguyen3, J. Leu4, E. F. Gillespie2, and R. Peng5; 1University of Washington - Seattle, Seattle, WA, 2Department of Radiation Oncology, University of Washington/Fred Hutchinson Cancer Center, Seattle, WA, 3Crozer-Chester Medical Center, Upland, PA, 4University of Washington School of Medicine, Seattle, WA, 5AAMC, Washington, DC, United States
Purpose/Objective(s): Commercial large language models (LLMs) are increasingly used for clinical trial screening, but their out-of-the-box performance remains uncertain. Synthetic data may help address privacy concerns and augment off-the-shelf LLM screening models. In the context of generating synthetic data to train a screening tool for the PREEMPT trial (NCT06745024), we conducted a blinded comparative study evaluating zero-, one-, and multi-shot prompting strategies to assess the realism of bone findings in synthetic CT reports.
Materials/Methods: Bone findings of CT reports were generated using GPT-4o mini under three prompting conditions (zero-, one-, and multi-shot) using anonymized examples. Prompt design was informed by review of reports from five practicing radiologists to capture realistic structural and stylistic elements.
A total of 150 reports (75 authentic and 75 synthetic, equally distributed across prompting conditions) were evaluated by six blinded raters (including four residents and one radiation oncologist) using a 4-point Likert scale assessing perceived authenticity. Ratings were dichotomized as authentic versus synthetic for analysis. One-way repeated-measures ANOVA was used to compare discrimination rate across prompting strategies.Results: Raters correctly identified 48% of all reports, indicating high clinical realism for the synthetic data. Rater discrimination rates for synthetic reports were low across all strategies: 0.336 (95% CI: [0.074,0.598]) for zero-shot, 0.240 (95% CI: [0.154,0.326]) for one-shot, and 0.248 (95% CI: [0.105,0.391]) for multi-shot. There was no statistically significant difference in performance between the three prompting conditions (p = 0.684), suggesting that the clinical realism of synthetically generated bone sections remained consistent regardless of prompt complexity. Fleiss' kappa (?) for synthetic report detection ranged from 0.04 to 0.079 across all groups, as raters consistently disagreed whether a certain report was authentic or synthetic.
Conclusion: Our findings indicate that raters were unable to distinguish synthetic from authentic reports better than chance regardless of prompting strategy. Synthetic data may serve as a viable, privacy-preserving resource for training screening tools such as in the PREEMPT trial.