Main Session
Sep 28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology

2531 - Retrieval-Augmented Generation for Automated Identification of Neoadjuvant Chemotherapy from Unstructured Clinical Text in Breast Cancer

10:45am - 12:00pm ET
Poster Hall - Exhibit Hall A
Screen: 18
POSTER

Presenter(s)

Hongyu Liu, PhD - Cedars Sinai Medical Center, Los Angeles, CA

H. Liu1, K. Tsang2, M. Kanthamneni3, N. Tatonetti2, and T. Dou4; 1Cedars-Sinai Medical Center, LOS ANGELES, CA, 2Department of Computational Biomedicine, Cedars-Sinai Medical Center, West Hollywood, CA, 3Department of Radiation Oncology, Cedars-Sinai Medical Center, West Hollywood, CA, 4Department of Radiation Oncology, Cedars-Sinai Medical Center, LOS ANGELES, CA

Purpose/Objective(s):

It is imperative to ascertain whether a patient diagnosed with breast cancer has undergone neoadjuvant chemotherapy (NAC), a systemic treatment administered prior to surgical resection, to ensure the accuracy of quality measurement, outcomes research, and clinical trial eligibility screening. Prior NLP methods have been used to extract cancer treatment information from clinical text, but they often struggle with the complexity and longitudinal nature of real-world narratives. In this study, we applied retrieval-augmented generation (RAG) to classify NAC cohort from longitudinal clinical narratives.

Materials/Methods:

Two RAG-based models were developed using pathology reports and clinical notes from a single-institution breast cancer cohort with manually curated labels. Model 1 (Keyword-Filtered RAG) retrieved neoadjuvant-related text chunks using 64 predefined terms, stored them in a vector database, selected the top 10 chunks per patient, and annotated negation status. Model 2 (Semantic RAG) removed keyword dependence by retrieving the top-K (3, 5, or 10) semantically similar chunks from 100-word document segments. Both models provided GPT-4o with structured prompts, source hierarchy (clinical notes and pathology reports), contradiction-handling rules, and indirect-evidence cues. Performance was evaluated with stratified splitting using accuracy, precision, recall, F1 score, and AUROC.

Results:

Model 1 attained an AUROC of 0.78. Model 2 achieved a stable AUROC of 0.94 with the 100-word chunk top-10 retrieval setting. Semantic retrieval captured cases with only indirect NAC evidence that keyword filtering missed. Contradiction-resolution prompting reduced false negatives from discordant documentation, particularly when pathology reports negated presurgical therapy while oncology notes confirmed NAC administration.

Conclusion:

We demonstrate that RAG-based classification of NAC receipt from unstructured clinical text is both feasible and effective. Semantic retrieval exhibited superior performance over keyword-dependent retrieval, achieving an AUROC that was 16 points higher. This approach presents the application of RAG to the identification of neoadjuvant therapy from EHR narratives.
Aspect

Variant

Approach

What Gets Indexed

Negation Handling

AUROC

Model 1

Keyword matching + RAG

Extract text windows around keyword matches (±500 chars)

Only chunks containing neoadjuvant 64 keywords

Explicit negation detection (50-char lookback for phrases like "no known", "did not receive") + metadata tag [NEGATED]/[AFFIRMATIVE]

0.7829

Model 2

100w * top3

Pure semantic similarity + RAG

Fixed word-count chunking (100 words, 20-word overlap)

All chunks from all documents (no keyword filter)

No explicit negation detection; relies on LLM prompt rules

0.8021

100w * top5

0.8750

100w * top10

0.9375