Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2493 - Accuracy of a Large Language Model in Automatic Extraction of Unstructured Clinical Data to Characterize the Pattern of Androgen Deprivation Therapy-Associated Hot Flashes

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 16
POSTER

Presenter(s)

Sonali Joshi, BS Headshot
Sonali Joshi, BS - California University of Science and Medicine, Colton, CA

S. Joshi1,2, J. Anzules1, Z. Huang1, A. Buckley1, Q. Feng1, M. Sayegh3, J. Y. C. Wong1, S. V. Dandapani1, S. M. Glaser1, S. Maroongroge1, C. J. Ladbury1, A. Chehrazi-Raffle4, T. B. Dorff5, A. Tam1, K. Man6, and Y. R. Li1; 1Department of Radiation Oncology, City of Hope National Medical Center, Duarte, CA, 2California University of Science and Medicine, Colton, CA, 3Creighton University, Phoenix, CA, 4Department of Medical Oncology & Therapeutics Research, City of Hope National Medical Center, Duarte, CA, 5City of Hope National Medical Center, Duarte, CA, 6Department of Applied AI & Data Science, City of Hope National Medical Center, Duarte, CA

Purpose/Objective(s): Androgen deprivation therapy (ADT) is an essential component in the management of advanced prostate cancer, but it is associated with substantial side effects such as hot flashes. Studies on hot flashes have been limited due to inconsistent and unstructured documentation in clinical notes, as this is not a routinely captured patient-reported outcome (PRO) in most hospital systems. With advances in artificial intelligence (AI), large language models (LLMs) can potentially overcome some of these challenges in extracting accurate information from medical charts. We evaluated an in-house-developed LLM framework that leverages state-of-the-art foundation models with an oncology-fine-tuned retrieval model for clinical note abstraction, as an approach to structuring symptom data from free text.

Materials/Methods: We randomly selected 100 patients from a previously assembled study cohort of prostate cancer patients who received treatment at our institution. After excluding patients who did not receive ADT, 73 patients were eligible for analysis. A 700-word prompt was given to the LLM, which was then used to compile the following information from clinical notes: (i) ADT regimen and treatment start/end dates and (ii) individual hot-flash mentions and contextual indicators of symptom persistence or resolution. An independent chart review was conducted separately to verify the abovementioned clinical information. Data from LLM and independent review were then compared to assess data accuracy.

Results: We found that our LLM accurately captured the ADT regimen, including medication names and the regimen dates, for 46 (63%) patients. Reasons for inaccurate information capture include incorrect dates (n = 23 [32%]), incorrect medications (n = 7 [10%]), or accuracy in one regimen but not all (n = 6 [8%]), with multiple inaccuracies in data pulled for some patients (n = 7 [10%]).The model had 100% recall of patients with reports of hot flashes when compared with those identified by manual review (n = 31 [42%]). Also with 100% accuracy, the model identified the 42 (58%) patients with no reports of hot flashes found on manual review.

Conclusion: This proof-of-concept demonstrates that LLM-based abstraction can enable scalable, clinically meaningful symptom capture using unstructured oncology notes and provides a framework extendable to other therapy-related toxicities. Despite its success, there are still ways to improve the model, including changes to the prompt, such as expanding drug-name coverage. Based on these promising findings, we anticipate that in the future, LLM-based approaches can be a particularly useful tool to capture PROs that are either not collected in a structured manner or were not initially recognized to be clinically significant.