We use cookies to enhance your site experience. This website stores data such as cookies to enable essential site functionality, as well as marketing, personalization and analytics. By continuing to use our website, you consent to cookies being used. Review our Privacy Policy for more information.
Zidu Xu, PhD - Mass General Brigham, Harvard Medical School, Boston, MA
Z. Xu1,2, J. Zeng3, S. Zhou4, Y. H. Chen5, A. Yaghoubi6, V. Goddla7, L. Lehmann8, E. Sharon9, D. E. Kozono10, A. Revette9, D. Dligach6, and D. S. Bitterman3; 1Department of Radiation Oncology, Brigham and Women’s Hospital/Dana-Farber Cancer Institute, Boston, MA, 2Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School, Boston, MA, 3Department of Radiation Oncology, Mass General Brigham/Dana-Farber Cancer Institute, Harvard Medical School, Boston, MA, 4Division of Computational Health Sciences, Department of Surgery, University of Minnesota, Minneapolis, MN, 5Department of Data Science, Dana-Farber Cancer Institute, Boston, MA, 6Loyola University, Chicago, IL, 7Harvard College, Cambridge, MA, 8Harvard Medical School, Boston, MA, 9Dana-Farber Cancer Institute, Boston, MA, 10Department of Radiation Oncology, Dana-Farber Cancer Institute/Brigham and Women’s Hospital, Harvard Medical School, Boston, MA
Automated safety and quality oversight of large language model chatbots for cancer clinical trial informed consent
Purpose/Objective(s):
There is significant interest in using large language model (LLM) chatbots to make complex oncology information accessible and interactive, but safety risks persist. LLM-as-a-judge (LAJ) methods could offer a scalable alternative to costly human ratings and provide automated guardrails. We evaluated the potential of LAJ to guardrail the safety and quality of chatbots for cancer clinical trial informed consent, a high-impact use case to address gaps in patient understanding of trial options.
Materials/Methods:
We analyzed 50 trial consent forms. Forty-five forms were used to create an 800-pair question-and-answer (QA) evaluation dataset. Five forms were used to create a development dataset for engineering prompting configurations. We developed a seven-criterion scoring rubric including three binary safety violation criteria (individualized medical advice, critical harm risk, overly persuasive language) and four 5-point Likert quality criteria (transparency, accuracy, usefulness, clarity). Six clinicians manually rated QAs as gold standards. We evaluated 5 LLMs across various prompting configurations (Table). Performance metrics were mean sensitivity for detecting safety violations and Spearman correlation (ρ) to measure Likert-scaled rating agreement, compared with clinician labels.
Results:
Dual-clinician agreement on a subset of 292 pairs showed a mean kappa of 0.75. The table details the optimal LAJ configurations across models and domains. Mean safety sensitivity was 0.88. Mean sensitivity peaked for individualized medical advice at 0.95, followed by overly persuasive language at 0.91 and critical harm risk at 0.78. Mean Spearman ρ was 0.78 for transparency, 0.69 for accuracy, 0.51 for usefulness, and 0.29 for clarity. The optimal configuration was GPT-5.2 using chain-of-thought (COT) prompting aligned with human scoring logic.
Conclusion:
LAJ methods demonstrated high sensitivity for detecting safety violations in clinical trial informed consent QA tasks. Performance was stronger for document-grounded transparency and accuracy than for subjective usefulness and clarity. Future work will optimize performance for scalable chatbot evaluation and clinical oversight.