Main Session
Sep
28
PQA 03 - Digital Health Innovation and Informatics, Patient Safety & Quality, and Radiation and Cancer Biology
2531 - Retrieval-Augmented Generation for Automated Identification of Neoadjuvant Chemotherapy from Unstructured Clinical Text in Breast Cancer
Presenter(s)
Hongyu Liu, PhD - Cedars Sinai Medical Center, Los Angeles, CA
H. Liu1, K. Tsang2, M. Kanthamneni3, N. Tatonetti2, and T. Dou4; 1Cedars-Sinai Medical Center, LOS ANGELES, CA, 2Department of Computational Biomedicine, Cedars-Sinai Medical Center, West Hollywood, CA, 3Department of Radiation Oncology, Cedars-Sinai Medical Center, West Hollywood, CA, 4Department of Radiation Oncology, Cedars-Sinai Medical Center, LOS ANGELES, CA
Purpose/Objective(s):
It is imperative to ascertain whether a patient diagnosed with breast cancer has undergone neoadjuvant chemotherapy (NAC), a systemic treatment administered prior to surgical resection, to ensure the accuracy of quality measurement, outcomes research, and clinical trial eligibility screening. Prior NLP methods have been used to extract cancer treatment information from clinical text, but they often struggle with the complexity and longitudinal nature of real-world narratives. In this study, we applied retrieval-augmented generation (RAG) to classify NAC cohort from longitudinal clinical narratives.Materials/Methods:
Two RAG-based models were developed using pathology reports and clinical notes from a single-institution breast cancer cohort with manually curated labels. Model 1 (Keyword-Filtered RAG) retrieved neoadjuvant-related text chunks using 64 predefined terms, stored them in a vector database, selected the top 10 chunks per patient, and annotated negation status. Model 2 (Semantic RAG) removed keyword dependence by retrieving the top-K (3, 5, or 10) semantically similar chunks from 100-word document segments. Both models provided GPT-4o with structured prompts, source hierarchy (clinical notes and pathology reports), contradiction-handling rules, and indirect-evidence cues. Performance was evaluated with stratified splitting using accuracy, precision, recall, F1 score, and AUROC.Results:
Model 1 attained an AUROC of 0.78. Model 2 achieved a stable AUROC of 0.94 with the 100-word chunk top-10 retrieval setting. Semantic retrieval captured cases with only indirect NAC evidence that keyword filtering missed. Contradiction-resolution prompting reduced false negatives from discordant documentation, particularly when pathology reports negated presurgical therapy while oncology notes confirmed NAC administration.Conclusion:
We demonstrate that RAG-based classification of NAC receipt from unstructured clinical text is both feasible and effective. Semantic retrieval exhibited superior performance over keyword-dependent retrieval, achieving an AUROC that was 16 points higher. This approach presents the application of RAG to the identification of neoadjuvant therapy from EHR narratives.| Aspect | Variant | Approach | What Gets Indexed | Negation Handling | AUROC |
| Model 1 | Keyword matching + RAG Extract text windows around keyword matches (±500 chars) | Only chunks containing neoadjuvant 64 keywords | Explicit negation detection (50-char lookback for phrases like "no known", "did not receive") + metadata tag [NEGATED]/[AFFIRMATIVE] | 0.7829 | |
| Model 2 | 100w * top3 | Pure semantic similarity + RAG Fixed word-count chunking (100 words, 20-word overlap) | All chunks from all documents (no keyword filter) | No explicit negation detection; relies on LLM prompt rules | 0.8021 |
| 100w * top5 | 0.8750 | ||||
| 100w * top10 | 0.9375 |