2423 - Reporting Completeness of Clinical Trials Evaluating Large Language Model Interventions in Medicine
Presenter(s)
D. Chen1, A. Blayne1, T. Ning1, I. Chattha1, and D. Bitterman2; 1University of Toronto, Toronto, ON, Canada, 2Massachusetts General Brigham, Boston, MA
Purpose/Objective(s): Large language models (LLMs) are increasingly tested in prospective clinical trials, but their generative outputs makes transparent, complete reporting essential for reproducibility and safety appraisal. This review mapped the landscape of the LLM interventions in clinical trials and quantified reporting completeness using reporting guidelines.
Materials/Methods: We conducted a systematic search on December 6, 2025, of Embase, MEDLINE, and CENTRAL for trial reports and protocols, and ClinicalTrials.gov and WHO ICTRP for trial registrations published from 2017 to present. Two reviewers independently screened and extracted trial and intervention characteristics. Three calibrated raters assessed item-level reporting using eligible reporting guidelines matched to study design and purpose, including TRIPOD-LLM, CONSORT 2025 and CONSORT-AI, SPIRIT 2025 and SPIRIT-AI, CHART, and the WHO Trial Registration Dataset. The study outcomes were overall study and guideline item-specific reporting completeness.
Results: The systematic search yielded 461 records, including 68 trial reports, 15 protocols, and 378 registrations. Trial reports were predominantly small sample-sized, single-site randomized studies published in 2024–2025, spanning diverse clinical specialties and primarily targeting patient-facing education and management recommendations. Mean reporting completeness was modest across TRIPOD-LLM (40%), CONSORT (46%), and CHART (52%).
Conclusion: Clinical trials of LLM interventions prioritize narrative framing and commentary over the methodological detail needed for replication, risk-of-bias assessment, and governance of model performance over time. Across reporting guidelines, study context, objectives, commentary about benefits and limitations, and disclosure statements were often reported. However, critical methodological elements for reproducibility, including data and intervention provenance, input pre-processing, model versioning and access modality, inference parameters, output post-processing, as well as protocol, data, and code access, were inconsistently reported. Harms planning, monitoring, and reporting were frequently incompletely reported, limiting the critical appraisal of the safety of LLM interventions in healthcare. Clinical trials of LLM interventions show modest adherence to established reporting guidelines, with persistent underreporting of LLM-specific methodological details, reproducibility artifacts (protocols, data, code), and harms reporting required to support reproducibility and confident appraisal of safety risks before their clinical translation.