Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2467 - Patient Ratings and Variability of Physician vs. ChatGPT Responses in Cancer Communication

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 16
POSTER

Presenter(s)

Ramesh Gopal, MD, PhD Headshot
Ramesh Gopal, MD, PhD - University of New Mexico Cancer Center, Albuquerque, NM

R. S. Gopal1, A. C. Segura2, B. Tawfik1, B. Liem3, T. M. Schroeder4, and D. Y. Lee1; 1University of New Mexico Comprehensive Cancer Center, Albuquerque, NM, 2Department of Biomedical Engineering, University of New Mexico School of Engineering, Albuquerque, NM, 3New Mexico Minority Underserved NCORP, Albquerque, NM, 4University of New Mexico, Albuquerque, NM

Purpose/Objective(s):

Physicians aspire to consistently communicate complex and contextual information to cancer patients. Artificial Intelligence (AI) platforms like ChatGPT (GPT) are alternative sources of information and advice which can provide personalized information in natural language. We compared blinded patient ratings of physician versus GPT responses across four common cancer-communication scenarios with regard to (1) overall preference and rating and (2) relative variability of ratings.

Materials/Methods:

We surveyed 51 adult female breast cancer patients undergoing radiation. Participants evaluated GPT versus physician-authored responses to a subset of four questions (Q1-4; total 104 paired comparisons): whether to 1) inform an adult daughter about a cancer diagnosis, 2) switch treatment centers, 3) quit working after a cancer diagnosis, and 4) available cancer prevention strategies. Each scenario included two blinded responses to the same question - one from GPT and the other from one of three physicians. Participants indicated their preferred response and rated each on a 7-point Likert scale with regard to being Useful (U) in making a decision, providing needed Information (I), and showing Sensitivity (S) to their feelings (score of 1 best, 7 worst). We assessed the proportion preferring GPT (binomial test vs 50%), within-subject differences in Likert scores for GPT and physician responses (Wilcoxon signed-rank test) and score variability (compared by Levene’s test, standard deviation (SD) reported).

Results:

Across 104 paired responses, patients preferred the answers from GPT over physicians 71 (68%) vs. 33 (32%), p <.001. GPT was significantly favored regarding work leave (85.7%, p = 0.00018) and cancer risk reduction (71.4%, p = 0.036), while for family disclosure (53%) and switching centers (59%) the preference for GPT was not statistically significant (p = 0.46). Across all questions, Likert scores showed small preferences for GPT over physician responses for U (1.85 vs 2.24, p = 0.008), I (1.98 vs 2.52, p = 0.001) and S (2.07 vs 2.32, p=0.118). Likert scores of physician responses were significantly more variable than those of GPT as measured by SD: U (1.431 vs 1.097, p<0.001), I (1.609 vs 1.099, p<0.001), S (1.556 vs 1.215, p<0.001). Greater variability in physician scores across U, I and S combined was significant in 3 out of 4 scenarios: Q1 1.571 vs 1.160, p=0.002; Q2 1.708 vs 1.477, p=0.013; Q3 1.566 vs 0.63, p<0.001; Q4 0.630 vs 0.940, p=0.247.

Conclusion:

Overall, patients preferred GPT responses over physician responses, driven primarily by two of four scenarios. Physician responses had greater variability in patient ratings than GPT responses. Our findings suggest that physicians could utilize GPT responses as a starting point to make their own communication more consistently preferable to patients.