Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2491 - A Large Language Model-Based Context-Aware, Patient-Specific Clinical Information Retrieval System for Automatically Extracting Clinical Variables In Patient Outcome Prediction

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 16
POSTER

Presenter(s)

Bowen Jing, PhD - Department of Radiation Oncology, UT Southwestern Medical Center, Dallas, TX

B. Jing1, B. Tortelli1, C. Moore1, H. Jiang2, R. Hannan1, and J. Wang1; 1Department of Radiation Oncology, University of Texas Southwestern Medical Center, Dallas, TX, 2Medical Artificial Intelligence and Automation (MAIA) Lab, Department of Radiation Oncology, UT Southwestern Medical Center, Dallas, TX

Purpose/Objective(s):

It is hypothesized that stereotactic ablative radiotherapy (SAbR) act synergistically with immunotherapy to improve cancer control. To assess this hypothesis our group was interested in conducting a large retrospective review, but the labor-intensive nature of chart review has been a challenge. Retrospective studies are increasingly used to assess clinical questions and construct prediction models to assist oncologists in developing personalized management plans. However, the significant labor and clinical expertise required to extract treatment variables from large volumes of unstructured clinical notes remains a significant obstacle. In this study we aimed to develop a large language model (LLM) based clinical information retrieval system capable of extracting desired variables from clinical notes.

Materials/Methods:

We developed a treatment information retrieval system consisting of two key modules: (1) a text data vectorization and context retrieval module, and (2) a large language model–based, context-aware information retrieval module. Several clinical variables were automatically extracted from clinical text notes for each patient such as earliest date of cancer diagnosis, prior surgery, toxicity of immunotherapy, and toxicity grade of immunotherapy. Clinical text notes—including radiology, oncology, medication, and hospitalization notes—were collected for 33 patients who underwent SAbR in the Department of Radiation Oncology. Automatically retrieved variables were compared with manually curated ground truth to assess system performance. Accuracy, precision, recall, and F1 score were calculated for each extracted variable when applicable.

Results:

The proposed context-aware LLM retrieval system achieved reasonable accuracy in extracting several key clinical variables from unstructured notes, with accuracies of 0.88 for date of diagnosis, 0.79 for prior surgery, 0.70 for immunotherapy toxicity (F1 = 0.72), and 0.67 for toxicity grade (F1 = 0.65), demonstrating the capability across numerical, textual, binary, and categorical data types.

Conclusion:

Our result shows that a context-aware retrieval system is capable of extracting key clinical variables from unstructured notes. With additional refinement and optimization, this approach has the potential to substantially reduce manual abstraction effort and enable scalable data extraction for retrospective studies.

Table 1. Performance of the proposed retrieval system

Entity retrieved

Date of diagnosis

Surgery procedure

Toxicity of immunotherapy

Toxicity Grade of immunotherapy

Data Type

Numerical

Text

Binary

Categorical

Accuracy

0.88

0.79

0.70

0.67

F1

0.72

0.65

Recall

0.72

0.67

Precision

0.72

0.64