Main Session
Sep
28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology
2678 - Comparative Readability, Understandability, and Accuracy Analysis of AI-Generated Patient Education Materials Across Eight Large Language Models
Presenter(s)
Jane Zheng, MS - Drexel University College of Medicine, Philadelphia, PA
J. Zheng1, and B. R. Page2; 1Drexel University College of Medicine, Philadelphia, PA, 2Johns Hopkins University Department of Radiation Oncology, Washington, DC
Purpose/Objective(s):
Low health literacy and the lack of readable patient education materials are significant barriers to treatment adherence among patients with head and neck cancer (HNC), contributing to compromised recovery and increased healthcare utilization. Large language models (LLMs) have emerged as potential tools for support, however, their comparative performance across platforms and model versions in radiation oncology remains unclear. We evaluated the accuracy and readability of multiple LLMs in generating patient-centered content on radiation therapy side effects for HNC patients.Materials/Methods:
8 LLMs including ChatGPT (4o, o3, 4.5), Claude, OpenEvidence, Grok, and Gemini 2.5 (Flash, Pro) were evaluated for their ability to generate patient education material for 3 HNC diagnoses: cT3N2M0 HPV+ tongue base, cT3N1M0 EBV+ nasopharynx, and T1N1M0 oral tongue squamous cell carcinomas. Readability was quantified using Flesch Reading Ease (FRE) and the SMOG Index, benchmarking against AMA, NIH, and NCI recommended literacy levels. Understandability was assessed using the Patient Education Materials Assessment Tool (PEMAT). Differences were analyzed using the Friedman test, with Wilcoxon signed-rank tests (Holm-adjusted) for post-hoc pairwise comparisons. Tool consistency was validated using Pearson correlation.Results:
Friedman test revealed significant differences in FRE scores across the 8 LLMs (X2(7) = 15.78, p = .027). ChatGPT 4o produced the most readable content (Mean = 74.40, SD = 1.40), whereas ChatGPT 4.5 produced the most complex text (Mean = 57.15, SD = 3.38). SMOG Index scores showed a trend toward significance (p = .057). ChatGPT 4o was the only model to successfully reach literacy target based on current health literacy guidelines (Mean = 8.49). Conversely, Gemini 2.5 Flash produced significantly more complex content (Mean = 10.61). A strong negative Pearson correlation (r = -.7865) between the FRE and SMOG indices validated the consistency of the measurement tools. While ChatGPT o3 achieved the highest average PEMAT understandability (89.1%), no statistically significant differences was found across models (p = .143). Post-hoc pairwise Wilcoxon signed-rank tests with Holm adjustments did not identify specific inter-model significance, likely due to the small sample size (N = 3 prompts).Conclusion:
LLMs show potential for automating patient education in HNC; specifically, ChatGPT 4o was most effective for meeting established health literacy guidelines. However, other models produced content at reading levels too advanced for the general public, which may exacerbate health disparities if implemented without modification. While LLMs assist in content generation, rigorous clinician oversight is necessary to ensure that simplifying language for readability does not compromise the clinical accuracy or patient safety.