Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2422 - Trial-Agent: An Agentic LLM System for Accurate Trial-Specific Question Answering during Informed Consent

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 14
POSTER

Presenter(s)

David Chen, MD - University of British Columbia, Vancouver, BC

D. Chen, R. Moscatel, E. Levy, and K. Quinn; University of Toronto, Toronto, ON, Canada

Purpose/Objective(s): Large language models (LLMs) could assist clinical trial coordinators in responding to trial participant queries during the informed consent process. We developed Trial-Agent, an agentic, multi-LLM system that generates accurate responses, audits and improves the accuracy of its own responses, and detects adversarial queries for human review.

Materials/Methods: Trial-Agent comprises (1) a Coordinator LLM that drafts responses using the trial knowledge base, (2) a Moderator LLM that rates factual accuracy (Likert 1–5 rubric) against the trial knowledge base and provides constructive feedback, and (3) an optional revise-and-rerate loop that prompts the Coordinator LLM to improve the accuracy of its response. We evaluated all Coordinator–Moderator pairings of OpenAI’s GPT-4o, GPT-5 Chat, and GPT-5 Thinking LLMs across two prospective clinical trials, DEFEND (telehealth-based exercise program for cancer patients) and CAN-SILENCE (prophylactic scopolamine butylbromide to reduce death rattle in dying patients). Trial-Agent was systematically benchmarked using 100 in-scope queries and 50 adversarial, out-of-scope queries for each trial. Primary study outcomes were the agreement of the Moderator LLM accuracy ratings with human trial coordinator accuracy ratings, the Moderator LLM-rated accuracy of the Coordinator LLMs draft response and post-revision response, and the proportion of blocked adversarial queries.

Results: The GPT-5 Chat Moderator accuracy ratings was comparable to human trial coordinator accuracy ratings for the DEFEND (human: 4.69, 95% CI 4.63–4.74; LLM: 4.63, 95% CI 4.58–4.69; p=0.142) and CAN-SILENCE (human: 4.83, 95% CI 4.76–4.89; LLM: 4.86, 95% CI 4.82–4.90; p=0.182). The most accurate LLM acting as the Coordinator responding to participant queries was the GPT-5 Chat LLM for the DEFEND trial (4.84, 95% CI 4.81–4.88) and the GPT-5 Thinking LLM for the CAN-SILENCE trial (4.94, 95% CI 4.92–4.96). The revise-and-rerate loop significantly improved inaccurate responses in both the DEFEND (original: 4.04, 95% CI 3.87–4.22; revised: 4.84, 95% CI 4.76–4.92; p<0.001) and CAN-SILENCE (original: 2.93, 95% CI 2.19-3.67; revised: 5.00, 95% CI 5.00–5.00; p-0.002) trials. Without prompt-based guardrails, all adversarial, out-of-scope queries bypassed the system (DEFEND: 0/150; CAN-SILENCE: 0/150). A general prompt instruction to block out-of-scope queries improved detection of adversarial queries (DEFEND: 142/150; CAN-SILENCE: 143/150), while category-defined guardrails consistently blocked all adversarial queries (DEFEND: 150/150; CAN-SILENCE: 150/150).

Conclusion: Trial-Agent is an agentic LLM system that enables scalable, trial-concordant responses to participant queries about clinical trials by combining trial-grounded retrieval augmented generation, LLM auditing of response accuracy with self-revision, and prompt-based guardrails to support accurate trial knowledge communication that may be useful in informed consent.